LLM Evaluation Linguist - Google
Evaluated Gemini-related large language model outputs for Indonesian language use cases to ensure linguistic quality and user-centered performance. Conducted response assessment with a focus on instruction adherence, relevance, and overall usability while using fact-checking to improve reliability. Supported prompt creation and iterative testing for model evaluation workflows, leveraging expertise in linguistics and AI quality assurance. • Evaluated AI responses for linguistic quality, instruction adherence, and relevance • Performed fact-checking and content verification to identify inaccuracies • Created and refined prompts for model evaluation, testing, and training workflows • Conducted linguistic review and quality feedback to improve model and dataset performance