AI Trainer (Independent Contractor) — AI data labeling and prompt/response evaluation work
Evaluated AI outputs by challenging LLM limitations and performing factuality checking on generated responses. Created domain-expertise prompts for medical use cases, iterating to improve response quality. Focused on ensuring outputs were accurate, reliable, and aligned with medical subject requirements. • Evaluate AI response quality and factuality • Challenge model limitations via adversarial testing • Write and refine prompts grounded in medicine • Produce ratings/feedback for downstream improvements