AI Model Evaluator (Non-Labeling) - Outlier
You train and evaluate AI models across coding, mathematical reasoning, and transcription tasks to improve overall LLM performance and response quality. You deliver annotated datasets and quality feedback while benchmarking model outputs against defined accuracy standards. You collaborate with cross-functional teams to align evaluations with technical requirements and domain expectations.• Train and evaluate LLM outputs for accuracy across technical domains• Provide annotated datasets and actionable feedback for model improvement• Benchmark responses against performance standards• Collaborate with cross-functional stakeholders to refine evaluation approach