AI Data Annotation Associate – LLM Evaluation
Evaluated and ranked large language model (LLM) outputs for accuracy, relevance, coherence, and factual correctness across diverse domains. Applied structured annotation guidelines to label data consistently and improve model performance and response quality. Reviewed responses to detect hallucinations, logical inconsistencies, and ambiguous outputs, supporting enhanced model reliability at scale. • LLM output ranking and comparative evaluation of multiple AI-generated responses • Consistent labeling using provided structured annotation guidelines • Quality checks for hallucinations, coherence, relevance, and factual correctness • Met strict turnaround time and productivity benchmarks while collaborating on evaluation criteria