AI Trainer / Core Contributor — Outlier AI & Remotask (LLM response evaluation, RLHF-style rating, and prompt testing).
Evaluated, graded, and ranked LLM-generated responses using strict rubrics for truthfulness, helpfulness, clarity, formatting, and safety for reinforcement-learning-style alignment. Leveraged medical and IT domain knowledge to fact-check and validate complex scientific, medical, and technical outputs against expected correctness and constraints. Built and applied prompt-based tests and multi-turn evaluation criteria to reduce hallucinations and ensure instruction following across long conversations. • Performed multi-turn conversational evaluations focused on context retention and formatting/tone adherence. • Crafted complex prompts to probe model boundaries and identify edge cases. • Assessed logical reasoning, constraint adherence, and factual accuracy for production-grade data delivery. • Used Outlier/Remotask workflow tooling for ongoing response review and ranking tasks.