AI/LLM Trainer (AI response evaluator) at Scale AI, Remote
Reviewed LLM outputs to score and label responses for accuracy, relevance, clarity, and adherence to project guidelines. Built and maintained labeled datasets by tagging and classifying text examples to teach the model correct interpretation and response. Performed consistency checks on annotated data to identify ambiguous labels and common error patterns and to ensure quality standards are met. • Assigned scores/labels to inform downstream model training • Tagged, classified, and organized text examples for dataset curation • Checked annotations for consistency, correctness, and ambiguity • Wrote actionable feedback on edge cases and undesirable outputs to refine rubrics and protocols