AI Evaluation & Medical Domain Expert (LLM response evaluation and clinical accuracy assessment)
Conducted evaluation of LLM-generated clinical responses for factual accuracy and clinical soundness. Assessed clinical reasoning quality, factual correctness, and logical consistency to refine output reliability for medical use. Supported development of an LLM-powered clinical decision-support workflow using prompt design and response testing. • Evaluated medical accuracy of model responses • Rated/assessed clinical reasoning and factual correctness • Ensured logical consistency across outputs • Performed iterative refinement based on evaluation results