Freelance AI Data Specialist (Self-employed / Various Platforms)
Reviewed and rated thousands of AI-generated responses for quality, factual accuracy, helpfulness, tone, and adherence to safety guidelines. Conducted prompt engineering review by identifying unclear instructions and recommending improvements to elicit better behavior from large language models. Performed quality checks for issues such as hallucinations, biases, and harmful content in model outputs. • Evaluation criteria: accuracy, relevance, helpfulness, tone, and safety compliance • Prompt + response quality improvement for LLM instruction following • Red-flag detection: hallucinations, bias, and harmful content • Dataset coverage: general knowledge, technical topics, support dialogues, creative writing