AI Data Analyst & Quality Evaluator at RWS
Conducted comprehensive evaluation of large language model outputs using rubric-based scoring frameworks. Assessed responses for accuracy, relevance, safety, and alignment with human preferences while refining rating decisions for edge cases. Provided actionable qualitative feedback to improve model behavior through better training data and evaluation methodology. • Applied rubric criteria to determine appropriate ratings for ambiguous and safety-sensitive scenarios • Developed and maintained annotation/evaluation guidelines to ensure cross-team consistency • Analyzed behavioral patterns and recommended improvements with machine learning engineers • Mentored junior evaluators and participated in quality calibration to uphold standards