Freelance AI Trainer & LLM Evaluator (Remote)
Freelance AI Trainer and LLM Evaluator performing quality assessment of AI-generated outputs for accuracy, reasoning quality, safety, helpfulness, and factual consistency. Conducted comparative judgment and ranking workflows, including pairwise comparison of model responses, to improve overall model output quality. Provided prompt and response review feedback across domains such as reasoning, writing, research, and customer interactions. • Evaluated text responses for factual consistency and safety. • Performed response ranking and pairwise comparisons. • Reviewed prompts and model outputs to refine annotation guidelines. • Reported edge cases with actionable feedback to improve AI performance.