For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
R

Ryan W.

Machine Learning & AI Evaluation Specialist

USA flagMckinney, Usa

Key Skills

Software

Scale AIScale AI

Top Subject Matter

LLM Output Evaluation
AI Reasoning
AI Evaluation

Top Data Types

TextText
ImageImage
Computer Code ProgrammingComputer Code Programming

Top Task Types

Text GenerationText Generation
Object DetectionObject Detection
Text SummarizationText Summarization
RLHFRLHF
Evaluation/RatingEvaluation/Rating
Computer Programming/CodingComputer Programming/Coding

Freelancer Overview

Machine Learning & AI Evaluation Specialist. Brings 11+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Scale AI, Internal, and Proprietary Tooling. Education includes Master of Science, Boston University (2014) and Bachelor of Science, Northeastern University (2010). AI-training focus includes data types such as Text and labeling workflows including Evaluation and Rating.

Labeling Experience

Senior Data Scientist, AI Evaluation

TextText

As a Senior Data Scientist at Google, I designed and executed structured evaluation workflows for AI-generated content. I focused on reasoning consistency, factual reliability, and benchmark validation for large-scale LLM applications. My responsibilities included collaborating on AI benchmarking and enhancing data quality standards in AI evaluation contexts. • Developed testing frameworks for evaluating LLM model outputs. • Implemented benchmark validation tasks to ensure high factual reliability. • Led efforts in model reliability and AI-generated response quality control. • Improved cross-team processes for human-in-the-loop and analytical quality assurance.

2019 - Present
Scale AI

Machine Learning & AI Evaluation Specialist

Scale AIScale AITextText

As a Machine Learning & AI Evaluation Specialist at Scale AI, I evaluated LLM-generated outputs for reasoning quality and factual consistency. My work involved performing rigorous error analysis and benchmark validation on AI-generated text responses. I contributed to scalable evaluation frameworks and quality assurance methodologies for AI training pipelines. • Analyzed large volumes of LLM output for structured logic and analytical correctness. • Identified and documented reasoning failures, supporting improvements in model reliability. • Conducted quality audits across extensive machine learning annotation systems. • Supported initiatives in AI safety, prompt review, and analytical response verification.

2016 - 2019

Education

B

Boston University

Master of Science, Data Science

Master of Science
2012 - 2014
N

Northeastern University

Bachelor of Science, Computer Science

Bachelor of Science
2006 - 2010

Work History

G

Google

Senior Data Scientist

Boston
2019 - Present
S

Scale AI

Machine Learning & AI Evaluation Specialist

Remote
2016 - 2019