LLM Evaluator & AI Trainer (Outlier, Remote)
Created SWE-bench-style engineering tasks to assess model performance on debugging, repository navigation, and code review exercises. Evaluated LLM-generated code for correctness, reasoning quality, edge-case handling, and compliance with specified project requirements. Designed structured validation workflows and automated test cases to ensure reproducibility and consistent software behavior. • Developed multi-step engineering tasks with clear acceptance criteria and expected outputs. • Performed annotation and benchmarking focused on instruction following and technical accuracy. • Built reproducible execution environments for consistent evaluations. • Focused validation on code quality, reliability, and software behavior.