Multilingual Quality Evaluator (Project-based) — AI Training / Annotation Work
Conducted structured bilingual (EN/ID) evaluation of AI model outputs against defined quality standards for insight, reasoning, and minimum-length requirements. Applied strict correctness criteria by flagging factual corruption (e.g., wrong dates/times) as major errors even when the response was otherwise fluent. Produced clear justifications tied to the rating rubric to enable iterative model improvement. • Rated responses for factual accuracy, reasoning depth, and instruction-following per rubric. • Reviewed outputs in both English and Indonesian using the same evaluation standards. • Documented why each rating was assigned to support downstream training adjustments. • Ensured corrupted or inconsistent facts were treated as critical failures.