LLM Evaluation and Text Generation – CrowdGen (Appen)
As part of the CrowdGen project by Appen, I worked as a LLM Evaluation and Text Generation Specialist, supporting the development and fine-tuning of a multilingual large language model. I labeled text data in English and Spanish by rating AI-generated responses based on fluency, coherence, relevance, and cultural appropriateness. I performed classification and ranking tasks, compared outputs for quality scoring, and created prompts for fine-tuning instruction-following models. I adhered to strict quality guidelines and met performance benchmarks, consistently delivering accurate and high-quality annotations.