RLHF Evaluator & Prompt Engineer
Conducted Reinforcement Learning from Human Feedback (RLHF) and text categorization for complex fintech automation workflows. Tasks included designing zero-shot and few-shot prompts, evaluating large language model (LLM) responses (ChatGPT, Claude, Gemini) side-by-side, and annotating outputs for factual accuracy, logical consistency, and safety. Strictly adhered to complex rating guidelines to produce high-fidelity training data.