Senior AI Training Expert & Code Evaluator (AI training and evaluator; RLHF preference tuning and dataset-quality validation).
Trained and evaluated AI preference models by applying RLHF methodologies to score responses for truthfulness and alignment. Authored structured multi-step prompt-and-answer materials to support LLM fine-tuning workflows. Performed systematic review of AI-generated code scripts to prevent execution edge cases.• Created and authored logically complex prompt/response datasets for model fine-tuning.• Evaluated model preference profiles using truthfulness and alignment scoring (via JSON/XML/Markdown).• Reviewed Python and SQL generated scripts for syntax and logical correctness.• Partnered with engineering leads to refine annotation guidelines and reduce dataset anomalies by 20%.