Freelance AI Training Specialist (Generalist)
Evaluated large language model (LLM) outputs for fluency, coherence, and natural tone. Assessed multi-turn conversations by annotating and ranking responses to improve handling of complex instructions and context. Performed quality checks on accuracy, relevance, and safety of generated text. • Fluency and coherence evaluation • Ranking and ordering conversation outputs • Accuracy/relevance/safety rating • Multi-turn prompt-response assessment