AI Data Annotator and Model Evaluator (Freelance) — LLM ranking, fact-checking, safety/compliance evaluation
Evaluated and ranked LLM outputs for correctness, helpfulness, harmlessness, and overall quality against safety and performance guidelines. Performed fact-checking by verifying claims against primary sources and assessing coherence and long-term reasoning quality. Produced qualitative feedback and engineering-ready evaluations to improve model behavior and reduce harmful or irrelevant responses. • Multi-turn dialogue assessment and evaluation • Red-team style testing for bias, harm, and unsafe behavior • Technical justification of model successes/failures • ETA (Expert Tier Analysis) labeling on outliers and complex reasoning/writing prompts