AI Evaluator—LLM Output Evaluation and Safety Testing
I participated in the evaluation of large language model (LLM) outputs for correctness, safety, and compliance with guidelines. My work involved ranking, reviewing, and annotating model responses in agentic task workflows. I focused on adversarial prompt design and model output assessment to support AI safety measures. • Reviewed and rated model outputs for factual accuracy and appropriateness. • Conducted RLHF annotation tasks on generated text to fine-tune LLM performance. • Designed and tested adversarial prompts to evaluate model resilience to unsafe behaviors. • Collaborated with AI researchers to assess and improve agentic workflows.