AI Trainer (Remote, 2023-Present)
Evaluated and ranked AI-generated responses based on accuracy, relevance, reasoning quality, and safety. Identified factual errors, logical inconsistencies, and hallucinations in model outputs to support model improvement. Wrote and refined clear, constraint-based prompts to test model behavior across multiple tasks. Applied rubric-based scoring and provided detailed justifications for ratings while adhering to strict task instructions and formatting rules. • Response ranking by quality and safety • Hallucination and inconsistency detection • Prompt writing for behavioral testing • Rubric scoring with justification and guideline compliance