Response Rating Evaluator, TELUS International
Evaluated and audited LLM responses for complex legal, regulatory, and financial compliance using strict multi-step grading rubrics. Generated sophisticated multi-turn adversarial prompts grounded in nuanced statutory logic to stress-test model reasoning and deductive assumptions. Applied pairwise comparison methods to rank outputs and improve specialized conversational AI ranking performance. • Conducted deep linguistic and logical analysis to identify language-use patterns impacting model refinement. • Collaborated with data science teams to align agentic behavior with accepted standards and first principles. • Identified hallucination risks and checked legal correctness, underlying assumptions, and deductive logic. • Used structured scoring criteria to support rigorous quality control in high-stakes compliance contexts.