Multilingual Conversational AI Evaluation Workflow
Designed and tested a multilingual conversational AI evaluation workflow focused on Southeast Asian language and cultural contexts. The project involved creating structured annotation and QA processes for conversational datasets containing English, Malay, Chinese, and code-switched dialogue samples. Tasks included: Response quality evaluation Intent classification RLHF-style response ranking Toxicity and safety review Data quality assurance The workflow was designed to support scalable human-in-the-loop AI evaluation processes for conversational AI and multilingual LLM applications. Quality measures included structured review guidelines, consistency checks, and manual QA verification across multilingual samples.