Software Engineer (Generative AI Evaluation) - Outlier
Contributed to RLHF-based evaluation and quality-check workflows to improve model response reliability and consistency. Generated and validated high-quality training tasks using AI tooling and agents to support evaluation pipeline improvement. Focused on ensuring accurate and useful outputs for downstream model assessment workflows, requiring strong attention to detail and familiarity with evaluation concepts. Worked with AI agents and developer tools to refine task generation processes and improve quality across repeated evaluations.• Supported RLHF evaluation workflows and quality checks. • Used AI tooling and agents to generate and validate training tasks. • Improved consistency in model evaluation pipeline outputs. • Applied evaluation quality standards and iteration for reliable task creation.