Chinese QA Specialist (AI Model Evaluation) at Outlier (Remote)
Performed AI model evaluation by reviewing LLM outputs for quality and policy compliance. Provided structured, actionable feedback to improve prompt design and increase reliability of responses. Identified edge cases, inconsistencies, and hallucinations to strengthen model robustness. • Rated/assessed factual accuracy of generated content. • Assessed reasoning quality, safety, and instruction adherence. • Wrote structured feedback for prompt and response improvements. • Flagged issues such as hallucinations and inconsistencies.