LLM Output Evaluation & Prompt Engineering
Worked as an AI Training Data Specialist for Turing, contributing to the development and quality improvement of large language models. Core responsibilities included evaluating and rating LLM-generated responses for accuracy, fluency, relevance, and instruction-following. Created and refined prompts designed to test and improve model reasoning and language capabilities. Produced paraphrased and rewritten text to expand training datasets, and classified content across multiple categories using detailed annotation rubrics. Maintained consistently high accuracy across large task volumes in a fully remote, flexible environment.