QA Automation Engineer (AI training, annotation, and LLM evaluation) — Amura Health
Evaluated and reviewed AI chatbot and LLM responses for quality, including relevance, correctness, completeness, clarity, grounding, consistency, and safety. Checked instruction-following against prompts, performed hallucination detection by verifying support from provided context, and validated RAG responses for source alignment. Created structured golden datasets and regression-test labeled examples with prompts, expected outputs, labels, and evaluation notes to support model and prompt iteration. • LLM and prompt-response evaluation with rubric-based quality scoring • Grounding/ground-truth verification for RAG-based chatbot answers • Safety and policy-aware checks including prompt injection and sensitive data leakage testing • Output comparison and response ranking to select the better answer