LLM Output Evaluator – Quality Assurance for Question Answering Systems
Served as the primary evaluator for large language model outputs in UPSC exam and medical question-answering AI projects. Identified factual inaccuracies, model hallucinations, and lack of source grounding by reviewing LLM text responses. Applied consistent criteria for output evaluation and provided structured feedback for iterative improvement. • Used tracing and prompt engineering to test model output reliability. • Evaluated groundedness of responses using source documents and retrieval mechanisms. • Developed and documented steps for systematic model assessment in high-stakes domains. • Worked with an internal review workflow and LangSmith for tracking evaluations.