AI Trainer / LLM Evaluation Specialist (Contractor) — Remote
As an AI Trainer / LLM Evaluation Specialist, Angela evaluated large language model outputs for quality, factual accuracy, instruction-following, and reasoning. She assessed text-based responses and also evaluated video-based outputs by reviewing them against task expectations and identifying failure modes. She produced comparative ranking decisions to support RLHF workflows and documented rationale for quality assessments. • Wrote prompts and evaluation rubrics to measure model performance • Performed comparative ranking of outputs for RLHF-style training • Identified hallucinations, weak logic, omissions, vague answers, and policy violations • Maintained consistency across high-volume evaluation tasks