AI Prompt Design & LLM Evaluation Lab (Personal Project)
Designed and stress-tested 60+ structured prompts across leading LLMs to assess output accuracy, coherence, and depth of reasoning for AI evaluation workflows. Produced structured evaluation reports documenting prompt behavior, model failure modes, and feedback guidance that are relevant to RLHF-style training and iteration. Practiced writing detailed feedback intended to support model fine-tuning and AI safety assessment in red-teaming contexts. • Prompt Engineering and LLM Benchmarking • RLHF feedback-oriented documentation and evaluation reporting • Adversarial prompting to surface model failure modes • Building a prompt library to reduce response iteration cycles