AI Model Evaluation & Prompt Engineering – NEUCLEUS Project
During the NEUCLEUS project, I rigorously tested and evaluated AI model outputs. I crafted intricate prompts to probe agent reasoning, rated model responses for both factual accuracy and safety, and debugged unexpected behavior to guide future training. This process closely mirrors real-world LLM evaluation and RLHF tasks. • Developed adversarial prompt scenarios for LLM/agent testing. • Provided structured ratings and qualitative feedback on model outputs. • Analyzed and reported response failures to inform system improvements. • Supported iterative enhancement of model safety and reliability.