AI Training Engineer (Contractor)
This role involved designing and authoring benchmark tasks for AI coding agents by leveraging real GitHub pull requests. I created multi-category grading rubrics to evaluate coding agent performance in areas such as functionality, robustness, code style, development trajectory, and toolchain usage. As part of the evaluation pipeline, I validated model output trajectories, inspected knowledge bases, and ensured justification fairness across various assessments. • Developed comprehensive and realistic programming benchmark scenarios. • Established detailed, inter-rater consistent grading rubrics for evaluation. • Validated and analyzed AI-generated code trajectories for fairness and correctness. • Collaborated remotely to optimize AI training and assessment methodologies.