Senior AI Data Annotator & LLM Evaluator (Remote, Contract; multiple AI platforms)
Responsible for evaluating thousands of LLM-generated responses for factual accuracy, coherence, tone appropriateness, and strict adherence to provided instructions across general knowledge, STEM, creative writing, and coding domains. Produced structured comparative rankings and written rationales to support RLHF-style feedback and iterative model improvement. Identified, documented, and categorized recurring failure modes such as hallucinations, instruction drift, and unsafe content using standardized evaluation rubrics. • Performed comparative response scoring, ranking, and rating based on rubric criteria • Conducted red-teaming and adversarial prompt testing to surface model vulnerabilities • Flagged toxicity and safety issues and reviewed output consistency against requirements • Calibrated guidelines with QA teams and resolved rubric edge-case disagreements