Data Scientist Intern: LLM Prompt Engineering and Evaluation
As a Data Scientist Intern, I developed and evaluated LLM workflows involving prompt engineering and LLM tuning for recommendation and extraction tasks. I built and automated pipelines for LLM evaluation using DeepEval, Uptrain, and Langfuse, focusing on trace retrieval, response quality scoring, and hallucination detection. I engineered both internal and external tools to improve AI model performance and reliability through data annotation and structured feedback loops. • Developed prompt/response datasets for LLM training, tuning, and evaluation • Automated data pipelines for LLM data extraction, annotation, and feedback aggregation • Implemented internal workflows to measure and monitor LLM outputs for relevance and quality • Utilized internal/proprietary tooling with Langfuse, DeepEval, Uptrain, and DSPy for AI workflow engineering.