AI Training Specialist / Data Evaluator (Human feedback for LLM, image, and video model evaluation)
Provided human feedback to evaluate AI outputs across language, image, and video models for correctness, safety, realism, and policy compliance. Designed and tested adversarial prompts to identify hallucinations and guardrail failures while comparing responses across scenarios to find edge-case errors. Maintained structured logs and guideline-driven quality tracking to support iterative improvements to AI model performance. • Evaluated model outputs for correctness and guideline adherence • Rated and explained findings to document failure modes and safety risks • Tested prompt variations to explore limits and behavioral weaknesses • Logged prompts, outputs, and evaluation results for reporting