Senior AI Response Evaluator & Language Quality Analyst (Independent Contractor | Remote)
Independently evaluated and annotated thousands of LLM-generated responses across general knowledge, reasoning, mathematics, policy, and everyday conversational scenarios. Labeled outputs using structured rubrics and taxonomies to score logical coherence, reasoning depth, completeness, tone, and user alignment. Performed source-based fact-checking to identify inaccuracies and unsupported claims, generating actionable feedback for model improvement. • Produced RLHF human feedback artifacts for reinforcement learning pipelines • Compared multiple model outputs and determined optimal responses via rubric-based evaluation • Maintained annotation consistency by following detailed benchmarks and evaluation guidelines • Delivered written, reproducible feedback aligned to taxonomy-driven scoring criteria