AI Data Labeler and LLM Evaluator
I contributed to LLM evaluation and data labeling projects focused on AI model output quality. My work involved designing evaluation frameworks, scoring rubrics, and comparative analysis of model-generated text. I applied statistical knowledge to assess response accuracy, consistency, safety, and relevance. • Performed structured evaluations and A/B testing of LLM outputs • Developed scoring systems and quality assurance protocols for annotated data • Conducted detailed annotation using both English and Chinese language skills • Ensured meticulous attention to annotation guidelines and consistency