Freelance AI Data & Japanese Language Evaluation Contributor
Provided guideline-based evaluations of Japanese AI-generated responses for fluency, naturalness, factuality, and instruction-following. Conducted side-by-side comparisons of model outputs and selected the stronger response based on quality and user intent. Identified issues such as hallucinations, mistranslation, incomplete answers, and poor reasoning while delivering written rationales for decisions. • Evaluated Japanese prompt/answer tone, grammar, cultural fit, and clarity • Performed model response ranking using project rubrics • Offered concise written feedback explaining evaluation choices • Worked independently on remote AI evaluation tasks requiring high attention to detail