Independent AI Engineer (LLM Output Evaluation & Annotation)
As an independent AI Engineer, I evaluated outputs from multiple Large Language Models (LLMs), including Claude, Qwen, and GPT-4, to improve performance. I developed internal rubrics and methodologies for judging adherence to instructions, factual accuracy, tone consistency, and hallucination detection. My work involved comparative assessments and prompt-based iterations to systematically enhance model responses. • Conducted bilingual (Mandarin/English) annotation and evaluation for RLHF and LLM output comparison tasks. • Created and applied original evaluation rubrics for LLM preference data and error analysis. • Executed prompt-quality annotation and side-by-side model judgment across multiple scenarios and user queries. • Documented prompt engineering process and published technical guides for the developer community.