Generalist AI Trainer (RLHF & LLM evaluation) | Scale AI (Remotask) | Contract
Provided RLHF, LLM evaluation, response ranking, and training data generation for large-scale AI model development pipelines using structured rubrics. Evaluated AI-generated responses for accuracy, tone, helpfulness, safety, and guideline compliance, including complex edge cases with consistent judgment. Completed preference ranking by comparing side-by-side model outputs and writing justifications aligned to defined quality criteria. • Compared model outputs and selected the better response using explicit RLHF quality criteria. • Performed binary classification, structured scoring, and categorization across customer service and conversational AI interactions. • Identified and documented hallucinations, misleading content, unsafe outputs, and policy violations for downstream decisions. • Contributed multilingual (English/Swahili) dataset support with culturally grounded context for diverse users.