Independent AI Evaluator & Prompt Engineer (2024–Present)
Conducted side-by-side benchmarking evaluations of multiple LLMs on professional architectural and legal question answering. Applied RLHF-style assessment criteria (Helpfulness, Honesty, Harmlessness) and identified logical inconsistencies and hallucination risks. Produced evaluation findings and refinements to guide model improvements and prompt changes. • Benchmarking across GPT-4, Claude 3.5, and Gemini • Rating and scoring of response quality using HHH • Detecting domain-specific AI hallucinations and logic errors • Feeding detailed feedback to drive iterative refinement