LLM Prompt and Response Evaluation Specialist
I executed manual annotation, quality rating, and prompt version comparisons for large language models across multiple platforms. My work focused on identifying hallucinations, verifying boundary cases, and providing structured evaluation of LLM responses. Dify's Log & Annotation tool was heavily used for bulk rating and systematic verification. • Designed input templates and boundary test cases for prompt evaluation. • Manually annotated and rated model outputs for hallucinations and risks. • Tracked hit rates and exported batch results for prompt version accuracy. • Compared and evaluated the performance of different LLMs on standard test sets.