Multilingual AI Response Evaluation – Outlier AI
Evaluated and ranked AI-generated responses across multilingual datasets in English and Spanish. Assessed outputs based on accuracy, relevance, coherence, reasoning quality, factual consistency, and compliance with user instructions. Applied detailed evaluation rubrics and project guidelines to identify high-quality responses and support model alignment objectives. Contributed to the improvement of large language models by providing structured human feedback, detecting hallucinations, identifying linguistic errors, and analyzing complex reasoning tasks. Maintained high standards of quality and consistency while working on large-scale AI evaluation projects in a remote environment.