AI Data Annotator / Preference Ranking & Prompt/Response Evaluation
Conducted preference-ranking style evaluation of multiple AI responses and tested prompts to determine which outputs better satisfy user requirements. Applied strong Chinese language understanding to assess content categorization, harmful/low-quality detection, sentiment or stance judgment, and summary quality. Produced guideline-aligned labeling rationales and error tagging to support high-quality data creation for model improvement. • Rated and compared responses for usefulness, safety, and alignment with intent • Classified user intent and content categories; assessed sentiment/stance and summary quality • Detected harmful or low-quality content and documented error rationales • Tagged errors and recommended prompt or response changes for iterative training