AI Data Trainer & Technical Annotator (RLHF, Instruction Evaluation, Output Ranking)
Provided RLHF and instruction-response evaluations to assess model output quality, accuracy, and adherence to instructions. Reviewed code-related tasks for correctness and flagged errors in technical content while maintaining consistent annotation quality. Performed model output ranking and preference labeling to support downstream training and improvement efforts. • Evaluated instruction-response pairs • Ranked outputs by quality/accuracy • Reviewed code correctness and technical accuracy • Flagged errors in technical content