AI Data Trainer & Labeler (Contract) | Remote
Performed rubric-based evaluation, annotation, and ranking of 1,500+ LLM responses to measure factual accuracy, logic, and safety. Audited STEM-focused outputs by checking mathematical accuracy and debugging code-related reasoning to reduce errors. Maintained a 98%+ quality score through strict compliance with granular, rapidly evolving platform guidelines. • LLM response ranking using complex rubrics • Manual review of factuality, logic, and safety • Verification of engineering proofs and code reasoning • Quality audits against evolving annotation standards