AI Engineer – MultiAgentic RAG Application and LLM Prompt Evaluation
Led the development of a MultiAgentic RAG application utilizing feedback-driven processes for AI model training and evaluation. Oversaw prompt engineering and iterative prompt refinement to enhance large language model (LLM) accuracy. Applied benchmarking and metric analysis on large-scale proposal data to improve system performance. • Integrated LangChain, LangGraph, Langfuse, and Locust for robust AI workflow execution. • Achieved 80% accuracy over 10,000 backtesting proposals, demonstrating impact on model reliability. • Enhanced LLM prompts and tested outputs for continual improvement of AI understanding and response. • Coordinated team efforts to follow best practices in scalability and security for data-driven AI tasks.