AI evaluation and AI-assisted documentation (LLM evaluator / AI trainer-style duties)
Performed ongoing AI output evaluation tasks by comparing responses from multiple LLMs and checking for factual accuracy, logical consistency, and instruction-following quality. Used predefined rubrics and a neutrality-focused approach to assess reasoning, detect hallucinations, and support safety/policy awareness requirements. Assisted with generating structured SOPs, process maps, and business documentation that follow complex implementation and quality guidelines. • Compared model responses to find inaccuracies and inconsistencies • Checked outputs for hallucinations and guideline violations • Maintained rubric-based scoring consistency and neutrality • Generated and refined structured documentation with AI assistance