AI Generalist (LLM grader/evaluator) at Mercor (Remote), June 2026–Present.
Evaluate AI-generated content outputs across two grading tracks covering content quality (factual accuracy, logical coherence, prompt adherence, depth, completeness) and aesthetic quality (structure, formatting, visual clarity, tone, and contextual presentation). Perform comparative analysis of competing responses to identical prompts and write structured justifications to identify stronger outputs despite surface-level thoroughness. Detect hallucinations, factual distortions, and plausible-sounding errors across business, operations, finance, and productivity domains while maintaining calibrated scoring standards throughout high-volume evaluation cycles. • Code-like and formula/workflow error checking for tools such as Excel, Google Sheets, and business intelligence platforms • RLHF reference generation and alignment-focused assessment of model outputs • Documentation of recurring failure patterns (formatting regressions, structural inconsistencies, domain gaps) and escalation for model improvement • Constraint/prompt framing effect analysis to surface behavioral patterns visible only via multi-prompt comparison