AI Data Annotator and LLM Evaluator, Mistral AI (Remote Contract via AfterQuery)
Rated and assessed LLM outputs for relevance, appropriateness, and correctness using multi-label scales and ranking tasks. Applied written explanations to flag ambiguous, inconsistent, and out-of-scope samples to improve labeling guidelines. Performed multilingual annotation support to widen cross-lingual training coverage across active LLM projects. • Relevancy scoring on 1–5 scales and pairwise ranking • Binary classification and content appropriateness review • English↔Swahili translation annotation, prompt rewriting, and concise summarization • Provided structured rationale to reduce edge-case escalations and improve agreement