Freelance Contractor AI/ML Engineer (RL Engineer) - Bespokelabs.ai
Designed and reviewed reinforcement learning training and evaluation tasks for frontier AI models across MLE-Bench, SoftwareBench, Terminal-Bench, SRE/Dunlin, MLOps, and Agentic AI benchmark environments. Built production-grade environments evaluating long-horizon reasoning, coding, debugging, machine learning experimentation, infrastructure management, and autonomous tool use. Developed reward functions, hidden validation checks, grading rubrics, and benchmark scenarios for post-training, reinforcement learning, and model capability evaluation. Progressed from Task Creator to Pod Lead, leading technical reviews, maintaining benchmark quality standards, and mentoring contributors. Analyzed model failures and reasoning breakdowns to identify capability gaps and improve reliability, robustness, and performance of AI systems. Areas include RLHF, RLAIF, AI evaluation, benchmarking, agent evaluation, post-training, long-horizon reasoning, tool use, coding agents, machine learning agents, and model alignment.