RLHF Trainer and Evaluator for LLMs (Pareto.AI)
I conducted RLHF training and evaluation for large language models such as Claude and ChatGPT, focusing on finance and investment workflows. Responsibilities included reviewing outputs, testing model reasoning, and ensuring quality standards were met. Complex financial models were used as the context for annotating, scoring, and improving model performance. • Evaluated LLM outputs for logic gaps, hallucinations, and reasoning quality. • Designed and tested edge financial scenarios to enhance reliability. • Used Python and PowerShell to validate Code Interpreter outcomes for financial tasks. • Spent over 400 hours on annotation, evaluation, and workflow documentation.