AI Model Evaluation & Data Annotation Specialist Contract / Project (2025 – Present)
Evaluated and graded 500+ LLM-generated responses for logical reasoning, factual accuracy, mathematical correctness, and safety guideline adherence. Performed structured review of outputs to ensure compliance with complex constraints and to identify issues affecting alignment and trust. Used scoring and rubric-based judgment to improve dataset quality and downstream model behavior. • Assessed reasoning correctness, factuality, and mathematical validity • Checked adherence to complex safety guidelines • Identified and reported subtle contradictions and semantic errors • Supported training alignment through iterative feedback