AI Evaluator & Prompt Engineer
Worked as an AI Evaluator and Prompt Engineer on Project Pontius at Alignerr, annotating and evaluating AI-generated finance-domain text. Evaluated outputs for factual accuracy, reasoning quality, numerical soundness, tone, and adherence to quantitative finance and investment banking protocols. Produced RLHF preference rankings and designed evaluation rubrics used to calibrate other annotators and set benchmarks. • Developed golden set response benchmarks for ground truth scoring. • Wrote detailed qualitative rationales for RLHF-based model training. • Covered subject matter spanning investment banking and quantitative finance. • Ensured annotation accuracy across complex financial scenarios.