Mercor Intelligence – Agentic AI / LLM Systems Research Engineer (Consulting Engagement)
Led evaluation and improvement of large-scale conversational AI systems for reasoning accuracy, code generation quality, and real-world engineering applicability. Defined and applied frameworks to assess LLM performance across algorithm design, debugging, and system design scenarios. Provided structured feedback and annotations to enhance model alignment, response clarity, and user-centric communication. • Conducted validation of model outputs through code execution, benchmarking, and best-practice comparisons. • Identified weaknesses in model reasoning and contributed to improving reliability, consistency, and factual accuracy. • Evaluated code quality for readability, maintainability, and algorithmic efficiency against production-grade standards. • Collaborated with research/engineering teams on next-generation agentic workflows and conversational behaviors.