AI Output Evaluator (Self-Training Project)
Regularly evaluated and tested outputs from AI tools on business and finance prompts, focusing on accuracy and reasoning quality. Ranked and annotated AI-generated responses for clarity, factuality, and domain expertise as part of personal RLHF and evaluation projects. Reviewed and verified AI-generated financial summaries for correctness, while building a comprehensive log of evaluated prompts and model behavior. • Performed LLM output evaluation across financial, accounting, and economics queries • Applied RLHF preference ranking to model outputs using internal logs • Used Outlier and Scale AI platform modules for hands-on practice • Developed expertise in identifying factual, reasoning, and hallucination errors