Freelance AI Content Trainer & Data Annotator (DataAnnotation.tech | Outlier | Alignerr)
Performed weekly evaluation and ranking of 500+ LLM-generated responses for instruction-following, factual accuracy, reasoning quality, and tone to support RLHF pipelines for frontier models. Engineered high-difficulty prompts across STEM, humanities, coding, and creative domains while meeting strict quality rubrics and NDA requirements. Conducted red-teaming and adversarial testing to identify safety violations, hallucinations, and bias patterns that informed model fine-tuning decisions. • LLM response evaluation and preference ranking • SFT training data creation including ideal responses and ranked preference pairs • Red teaming/adversarial testing and quality scoring • Asynchronous collaboration with QA leads to update annotation guidelines