AI Evaluation Specialist / Prompt Engineer (RLHF & SFT Specialist)
Performed RLHF and SFT evaluation work by ranking LLM outputs for technical accuracy, helpfulness, and safety. Authored high-quality ground-truth style responses to model training with consistent instruction-following and logical reasoning. Detected hallucinations and provided written justifications for model corrections to improve reliability across STEM, coding, and design domains. • Rated/ranked candidate model responses against safety and correctness criteria • Wrote and validated “golden response” ground truths for instruction-following • Identified subtle factual and logical inconsistencies and documented correction rationales • Designed multi-turn prompt scenarios to test edge cases and conflicting constraints