Data Labeling Specialist - Model Response Evaluation
In this role, I evaluated AI model responses through helpfulness, accuracy, instruction following, completeness, and language quality criteria. I performed bilingual (Chinese and English) review for both model outputs and guideline interpretation. My work contributed to ongoing improvements in AI model reliability and training feedback cycles. • Scored model-generated answers for various quality metrics • Checked responses for accuracy, completeness, and adherence to instructions • Provided constructive commentary and suggested improvements • Carefully documented ambiguous or problematic examples for further review