AI Evaluation Specialist (Freelance, Remote)
Evaluated AI-generated code across multiple languages and domains for correctness, quality, security, and adherence to engineering standards. Built QA frameworks and test cases to systematically assess LLM and code outputs against expected behaviors and constraints. Validated model responses against defined criteria and provided actionable feedback to improve model reliability. • Correctness validation • Quality and security review • LLM test suite and behavioral specification design • Actionable feedback for improving reliability