AI Engineer Intern - LLM Response Evaluation (ArchBlock)
Developed and implemented an evaluation framework for LLM-generated responses, utilizing both human scoring and automated NLP metrics. Evaluated over 1,000 text-based question-answer pairs to identify system weaknesses in tabular context retrieval. Collaborated with other engineers to optimize the RAG pipeline using insights from this structured evaluation. • Built multilayer evaluation rubric integrating human-in-the-loop scoring and automated LLMs-as-judges• Assessed context completeness, response accuracy, and relevance in financial document Q&A• Employed OpenAI and Anthropic batch API evaluation for standardized human and AI judgment• Provided actionable recommendations for improving data retrieval and context engineering methods