AI Text Annotator & LLM Output Evaluator (Academic & Project-Based, Self-Directed)
Performed rubric-based evaluation and scoring of AI/ML model outputs to identify labeling errors, hallucinations, reasoning inconsistencies, and classification mismatches. Applied annotation guidelines to validate 100,000+ rows while maintaining consistency and full audit traceability across multi-reviewer workflows. Leveraged multilingual capability (native Arabic and fluent English) to assess contextual correctness, semantic errors, and cross-lingual accuracy in multilingual NLP outputs. • Risk-signal label and sentiment classification schema design for human-in-the-loop checks before stakeholder delivery. • Compared predictions against ground-truth labels to find systematic errors and edge cases. • Documented discrepancies in structured reports to guide model retraining cycles. • Conducted multi-reviewer quality assurance with structured feedback labeling and traceability.