AI Evaluation Specialist, Outlier.ai
Evaluate and rank AI-generated responses in Arabic and English using truthfulness, helpfulness, clarity, and safety criteria. Develop and author nuanced Arabic prompts (MSA and regional dialects) to stress-test LLM boundaries and capabilities. Provide detailed English written justifications that explain superiority or logical flaws in candidate responses. • Truthfulness and factual accuracy checks against reliable sources • Safety- and guideline-based evaluation of outputs • Prompt writing for targeted LLM behavior testing • Written rationale and comparisons for evaluation decisions