Freelance AI Trainer / LLM Evaluator / Bilingual Data Annotator
Evaluated AI-generated explanations, essays, code answers, and reasoning responses in both English and Mandarin for factual accuracy, helpfulness, and adherence to instructions. Developed and applied evaluation rubrics for student writing, AP psychology, admissions emails, and technical debugging tasks. Rewrote weak or hallucination-prone outputs to ensure clearer, safer, and more human-readable responses. • Focused on improving clarity, instructional alignment, and tone control. • Identified and documented ambiguous logic, privacy risks, and low-quality reasoning. • Bilingual assessment of model accuracy on technical and general content. • Stress-tested AI outputs for hallucinations or unsupported claims.