AI Evaluator & Data Annotator (Contractor) - LLM evaluation and code QA (Dec 2025 – Present)
Performed rigorous testing and quality assurance for Large Language Models to improve reasoning and Python code generation. Evaluated model outputs for logical consistency, factual accuracy, and Python code functionality as part of an AI evaluation workflow. Produced structured feedback describing observed system behaviors to support engineering iteration and alignment/safety goals. • Assessed output correctness and reliability • Debugged and validated Python code generated by the model • Checked consistency and factuality of responses • Authored documentation for engineering teams