Data Annotator / LLM Response Evaluator (Independent Contractor, Remote)
Worked as an LLM data annotator and response evaluator for RLHF-style optimization by scoring outputs against multiple task criteria and safety requirements. Created and calibrated high-fidelity prompt training sets and detailed 20-point technical evaluation rubrics to ensure response quality and logical consistency. Audited model outputs against binary guidelines and used verification techniques to isolate errors such as factual issues and hallucinations. • Scored responses on Instruction Following, Completeness, Relevance, Accuracy, and Safety • Generated prompt training sets and 20-point technical evaluation rubrics for calibration • Audited outputs to identify hallucinations, bias, punts, and refusal-to-answer failures • Performed comparative analysis and provided preference justifications for response ranking