Freelance AI Data Annotator | Various AI Labs (Remote)
Provided RLHF ratings and comparative evaluations for LLM outputs across coding, reasoning, and instruction-following tasks. Annotated over 25,000 text samples covering sentence boundaries, coreference chains, toxicity classifications, and factual accuracy judgments for NLP research. Reviewed and red-teamed AI-generated code to identify logical errors, security issues, and hallucinations, then suggested corrected alternatives with rationale. • RLHF (Reinforcement Learning from Human Feedback) ratings and preference/comparative evaluation • Annotation for NLP tasks including boundaries, coreference, toxicity, and factuality • Red-teaming and corrected-code feedback with explanatory rationale • Calibration-guideline adherence with >95% agreement on checks