AI Financial Evaluator & RLHF Specialist (Freelance/Contract)
Freelance AI Financial Evaluator and RLHF specialist assessing and ranking LLM outputs for financial instruction-following and reasoning quality. Tasks included side-by-side comparisons, rubric calibration, and documenting failure modes to improve RLHF pipelines and model behavior in finance-specific contexts. Annotation work emphasized helpfulness, factual correctness, mathematical accuracy, and adherence to prompts while maintaining high throughput. • Evaluated and ranked LLM responses for earnings analysis, portfolio recommendations, market predictions, and macroeconomic reasoning. • Authored and refined financial prompts to probe derivative pricing, credit risk, and fiscal policy reasoning. • Performed SxS evaluations scoring helpfulness, factuality, math correctness, and instruction adherence. • Calibrated evaluation rubrics and flagged hallucinated figures, miscalculated ratios, and erroneous regulatory citations.