AI Coding Evaluator / Chinese LLM Evaluator
I evaluated and rated AI-generated code outputs from various language models in Chinese, focusing on correctness, maintainability, and engineering quality. My work included RLHF preference ranking, Chinese LLM output evaluation, and assessment of agentic multi-step coding tasks. I regularly reviewed AI output through a real-world software engineering lens to ensure reliability and task fitness. • Reviewed and rated Chinese code generation and RLHF outputs. • Evaluated technical Q&A, code explanations, and technical documents for quality. • Compared and judged outputs from multiple models and AI agents. • Provided detailed rationales and rankings in Chinese for model preference data collection.