AI Logic & Safety Evaluator / Reviewer — Mercor, Verity Labs (Remote)
AI Logic & Safety Evaluator/Reviewer at Mercor and Verity Labs assessed and scored LLM outputs using detailed rubrics. The work included reviewing other annotators’ outputs and running evaluations on logic, red-teaming, knowledge-base, safety, humanization, and creative/humor tasks. It also involved annotating text-based training data and correcting model outputs for accuracy and tone while complying with NDA and data-security requirements. • Designed and evaluated logic tasks in mathematical and natural-language formats for training and benchmarking • Conducted red-teaming and adversarial testing to identify safety failures and edge cases • Scored multi-turn conversation outputs for reasoning, coherence, alignment, logic, cultural context, safety, and instruction-following • Annotated and corrected text-based training data and model outputs for accuracy and tone