Freelance AI Data Trainer / LLM Evaluation Contributor
Worked as a freelance AI Data Trainer, evaluating AI model responses for accuracy, helpfulness, safety, and instruction-following. Compared multiple LLM-generated outputs, performed RLHF ranking, and identified linguistic issues including hallucinations and incorrect formats. Used GPT tools to enhance the evaluation workflow and ensure detailed, high-quality feedback. • Side-by-side LLM response and RLHF ranking tasks • Prompt writing and evaluation for varied domains (tourism, business, translation, etc.) • Issue detection for hallucinations, irrelevance, or factual errors • Workflow managed via GPT tools and collaborative documents