Golden Arena behavioural evaluation harness
Comparative behavioural scoring of AI models through five adversarial games and a Behavioural Index.
Hire this AI Trainer
Sign in or create an account to invite AI Trainers to your job.
I evaluate AI output and build the harnesses that do it at scale. On a paid contract for a TikTok-affiliated client I built the task set and scored 35 benchmark cases against a defined rubric, delivered across two approved milestones. I built Golden Arena, a five-game harness with a behavioural index that ranks models on how they actually play rather than on what they claim. Before an LLM document-extraction pipeline went anywhere near real accounting books, I ran a 25-invoice reliability battery against it and held the release until it passed. I work to written rubrics and stop-conditions, and I report an inconclusive result as inconclusive. I am native in English and French, so I can evaluate and compare output in both. My background is project-led rather than academic: two years building production TypeScript, Node and PostgreSQL systems, with open-source tooling published to npm and the official Model Context Protocol registry.
Comparative behavioural scoring of AI models through five adversarial games and a Behavioural Index.
Evaluated document-extraction outputs across 25 invoice cases for reliability and correctness.
Scored 35 benchmark cases against a defined rubric as part of an AI evaluation/crowdtesting contract.
HubSpot Revenue Operations Certified
HubSpot Digital Marketing Certified
AI Evaluation Contractor and Founder
Self-directed engineering practice