AI Output Rater / Model Evaluation Specialist — Feather
Provided human-feedback evaluation for frontier-model training using detailed HL grading rubrics. Rated AI-generated slide-deck artifacts for factual accuracy, instruction-following, design quality, and structural coherence to generate reward-relevant signals. Applied quantitative and technical domain knowledge to assess complex multi-step outputs where general raters may lack correctness context. • Evaluated outputs for the “Les Artistes — Artifacts” HL Grading campaign • Produced high-precision, granular feedback consistent across large batches • Graded technical content, document structure, and reasoning depth • Maintained rubric adherence for training pipeline inputs