LLM Evaluation Engineer (Freelance) - Outlier AI (Scale AI)
Served as an LLM Training Engineer for evaluating and improving AI coding tasks within RLHF-based training workflows. Conducted hands-on analysis of complex frontend, backend, full-stack, and scripting outputs to improve task quality and correctness. Applied prompt engineering, rubric-based evaluation, and code review practices to ensure instruction adherence and robust edge-case handling. • Evaluated 25+ complex AI coding tasks across multiple workflow categories to support RLHF LLM training pipelines. • Engineered prompt-response pairs with 40+ test cases and 20+ rubric checks per task to validate correctness and compliance. • Performed 90+ hours of independent QA on AI-generated code, identifying logical inconsistencies and requirement gaps. • Utilized JavaScript and React.js knowledge to assess and improve AI-generated solutions and evaluation artifacts.