LLM Red Team Medical Hallucination Validator
I designed and executed LLM red team test cases to validate model reliability in interdisciplinary medical, biological, and engineering contexts. The tasks included identifying and labeling hallucinations, logic errors, and domain-specific inconsistencies in LLM outputs. Comprehensive prompt refinement and output evaluation ensured robust assessment of LLM responses. • Tested and labeled LLM medical outputs for hallucinations. • Developed and applied targeted prompts and constraints. • Assessed reasoning, terminology, and causal interpretations. • Focused on specialized medical scenarios in LLM response analysis.