AI Training Specialist (Contract) | Scale AI
I evaluated outputs produced by Large Language Models for mathematical and logical correctness. My work included debugging AI-generated code snippets and verifying the integrity of multi-step problem solutions in Python and C++. I ranked and assessed model responses based on accuracy, helpfulness, and technical rigor. • Conducted Reinforcement Learning from Human Feedback (RLHF) on LLM outputs • Reviewed complex proofs and code for correctness and security • Used internal Scale AI platforms for annotation and ranking • Contributed to improving AI model reasoning and accuracy