AI Trainer / AI Evaluator / Reasoning Specialist (International Artificial Intelligence Projects, Remote)
Trained and evaluated advanced generative large language models (LLMs) by creating tasks, benchmarks, and evaluation prompts to measure reasoning quality. Applied reinforcement learning from human feedback (RLHF) workflows to optimize model behavior, accuracy, and safety. Produced and validated domain-specific datasets and ratings for complex math and physics reasoning. • Designed mathematical reasoning tasks across secondary to university-level topics • Evaluated physics/scientific reasoning for factuality, consistency, and response quality • Performed RLHF tasks in human-in-the-loop improvement pipelines • Built specialized travel-planning and reservation/booking agent datasets and prompts