AI Trainer / AI Content Contributor (Freelance) — prompt design, LLM response evaluation, and ground-truth rewriting
Trained and supported LLM behavior by rewriting sub-optimal model outputs into gold-standard ground truth to improve reinforcement learning (RLHF) outcomes. Evaluated and ranked model responses using criteria such as accuracy, truthfulness, tone, formatting, and safety guideline compliance. Produced detailed, objective justifications identifying logical or factual errors to explain why one response is preferred over another. • Crafted complex, high-quality prompts across multiple domains to test LLM boundaries. • Performed response comparison and ranking for quality, correctness, and adherence to requirements. • Wrote correction rationales citing specific logical or factual issues. • Generated improved reference responses for model reinforcement learning reinforcement.