AI Trainer / AI Model Evaluator — Outlier AI (Remote)
Evaluated and improved AI-generated responses for coding, reasoning, and technical problem-solving tasks. Designed prompts, rubrics, and evaluation frameworks to benchmark large language models on software engineering and mathematics topics. Reviewed model-generated code for correctness, instruction following, edge-case handling, code quality, and robustness. • Assessed responses using structured criteria (rubrics/frameworks). • Performed error analysis and failure identification to guide training. • Contributed to prompt engineering and response ranking workflows. • Focused on quality analysis of LLM outputs for technical domains.