LLM Response Evaluation & Text Annotation Specialist
Worked on AI training and text annotation projects focused on evaluating and improving large language model (LLM) responses for conversational AI systems. Responsibilities included reviewing AI-generated outputs for factual accuracy, instruction-following, grammar, tone, relevance, and safety compliance across a wide range of prompts and subject areas. Tasks also involved ranking multiple AI responses, identifying hallucinations and misleading information, labeling harmful or low-quality content, and applying detailed annotation guidelines to maintain consistency and reliability in training datasets. Contributed to prompt evaluation and quality assurance workflows designed to improve model performance and user experience. Reviewed structured and unstructured text datasets, categorized conversational content, and assisted in refining datasets for AI model fine-tuning and reinforcement learning tasks. Maintained high accuracy standards while working independently in remote environments and consistently meeting project deadlines and quality benchmarks.