Mechanic Glen V2 – Mathematical Response Evaluation and Rubric Design
Evaluated and ranked multiple AI-generated mathematical responses according to correctness, reasoning quality, instruction following, clarity, and mathematical rigor. Tasks involved comparing competing model outputs, identifying logical errors, verifying derivations, and producing detailed justifications for rankings. Additionally, the project required designing novel evaluation rubrics and grading dimensions to measure mathematical intelligence, reasoning quality, and problem-solving ability. Particular emphasis was placed on proof verification, mathematical accuracy, and rigorous assessment of model performance.