AI Trainer | AI Response Evaluator | Prompt Engineering & Financial Market Research (comparative AI response evaluation and fact-checking)
Conducted comparative evaluations of AI-generated responses from ChatGPT, Claude, and Gemini to assess factual accuracy, reasoning quality, clarity, depth, and neutrality. Reviewed responses for hallucinations and corrected inaccurate information related to financial markets, macroeconomics, and geopolitics. Applied structured reasoning and prompt engineering practices to improve the quality and reliability of LLM outputs. • Compared outputs across multiple LLMs (ChatGPT, Claude, Gemini) • Evaluated accuracy, reasoning quality, clarity, and usefulness • Detected hallucinations and corrected misinformation • Used prompt engineering to enhance response quality