Sama — AI Model Evaluator & Code Specialist
Served as an AI model evaluator and code specialist performing rubric-based review of AI-generated responses to assess correctness, quality, clarity, factuality, and edge-case handling. Conducted annotation review and dataset validation by identifying hallucinations, citation errors, factual issues, and bias patterns within evaluation outputs. Built repeatable evaluation workflows to maintain high reliability across repeated task types and improve labeling consistency for downstream teams. • Completed 500+ AI evaluation and annotation tasks covering code, research summaries, and long-form generated responses. • Maintained 95%+ inter-annotator agreement to ensure consistent model evaluation and labeling outcomes. • Applied APA/MLA editorial and citation standards during review of research/manuscript-style outputs. • Cleaned and reorganized large JSON datasets and developed a Python evaluation harness to improve review consistency.