AI Model Evaluator & Prompt Engineer
As an AI Model Evaluator & Prompt Engineer at Outlier AI, Alignerr, and Mercor, I performed rubric-based evaluation of language models and multimodal AI systems. My work included assessing outputs for clarity, correctness, and alignment, as well as conducting red teaming and adversarial testing to improve robustness. I also crafted prompts and detailed annotations to fine-tune and align next-generation AI systems.• Evaluated LLMs, agentic tools, and multimodal models with structured rubrics • Reviewed and scored model responses for multiple quality dimensions • Conducted dataset testing for marketing-related training flows • Wrote high-quality prompts and detailed annotation for fine-tuning pipelines