AI Data Evaluator — Telus Global (Remote)
Evaluated AI agent responses using detailed guidelines and scoring rubrics to assess accuracy, relevance, clarity, and safety. Identified hallucinations, bias, misinterpretations, and unsafe outputs across multiple domains through systematic review. Performed prompt testing to induce edge cases and improve alignment with instructions and user intent. • Compared multiple model outputs side-by-side and selected the best responses • Wrote rationales documenting scoring decisions and quality findings • Refined prompts to analyze model behavior and improve response alignment • Maintained accuracy and consistency across high-volume evaluations while following confidentiality protocols