AI Workflow Automation Engineer - LLM Response Evaluation
As an AI Workflow Automation Engineer, I performed structured evaluation of LLM responses across multiple open-source and proprietary models. This work involved systematically rating output for accuracy, coherence, safety, and adherence to prompts in both Chinese and English. I developed automated scripts and workflows to enhance consistency and accuracy in evaluation tasks. • Conducted head-to-head comparisons across 10+ LLMs. • Developed rubrics for instruction-following and response relevance. • Utilized Python and in-house tools to semi-automate review processes. • Focused on edge-case identification and documentation.