AI Trainer at Airtm (AI evaluation and annotation of model responses using complex grading guidelines).
Provided AI model evaluations to assess response accuracy, grammar, safety, and helpfulness against expected grading criteria. Applied hallucination detection and mitigation techniques by reviewing outputs and flagging unreliable or unsafe content. Delivered actionable human feedback to improve model behavior and reduce errors over iterative reviews. • Evaluated more than 10 AI model responses • Assessed accuracy, grammar, safety, and helpfulness • Supported hallucination detection and mitigation workflows • Used detailed feedback to inform RLHF-style training improvements