Technical AI Data Reviewer & Annotator (Remote)
Evaluated and benchmarked AI model outputs for technical accuracy, logical flow, and code efficiency across diverse programming prompts. Conducted high-complexity data annotation by ranking responses using RLHF-style paradigms to align outputs with user intent and safety guidelines. Authored multi-step technical rationales diagnosing why specific code generations succeeded or failed, including identification of subtle issues and factual errors. • Assessed logical correctness, edge cases, and syntax or semantic bugs in model-generated code • Ranked/selected preferred responses according to alignment with intent and safety guidelines • Produced structured justifications explaining failures such as hallucinations or logical fallacies • Performed quality assurance via detailed verification of reasoning steps and outputs