For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
W
Williams D.

Williams D.

Data Analytics & AI Evaluation Contributor (Remote)

USA flagN/A, Usa

Key Skills

Software

Other
ClickworkerClickworker
CVATCVAT
LabelboxLabelbox
OneFormaOneForma
MindriftMindrift
RemotasksRemotasks
SuperAnnotateSuperAnnotate
Scale AIScale AI
TolokaToloka
TelusTelus

Top Subject Matter

LLM alignment evaluation
RLHF preference ranking
safety/helpfulness/factuality rubrics

Top Data Types

TextText
ImageImage
AudioAudio

Top Task Types

RLHFRLHF
ClassificationClassification
Bounding BoxBounding Box
Text SummarizationText Summarization
Evaluation/RatingEvaluation/Rating
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
TranscriptionTranscription
Entity (NER) ClassificationEntity (NER) Classification

Freelancer Overview

Data Analytics & AI Evaluation Contributor (Remote). Brings 4+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Other. Education includes Bachelor of Science, University of Florida (2022) and Certificate in User Experience Design, Coursera. AI-training focus includes data types such as Text and labeling workflows including RLHF, Evaluation, and Rating.

Labeling Experience

AI Evaluation and Data Quality Contributor - Various Platforms

TextTextRLHFRLHF

You contributed to AI evaluation and dataset cleaning activities from a remote setting, focusing on preparing structured inputs and validating quality prior to engineering handoff. You reviewed outputs against safety, helpfulness, and factual accuracy rubrics and escalated edge cases through structured workflows. You also produced performance reporting on annotation throughput and accuracy trends using spreadsheet analytics and dashboard techniques. • Cleaned and validated structured datasets by removing duplicates, fixing schema issues, and verifying formatting • Performed rubric-based QA on AI-generated outputs and managed escalations for policy violations • Built weekly accuracy and throughput reports using Excel and Google Sheets • Supported RLHF pipeline comparisons using A/B-style evaluation workflows

2023 - Present

Data Analytics & AI Evaluation Contributor (Remote)

OtherTextTextRLHFRLHF

Labeled and quality-checked large NLP/LLM datasets using rubric-based frameworks to support reinforcement learning from human feedback pipelines. Performed human preference ranking and A/B-style response comparisons to contribute to model alignment iterations. Reviewed AI-generated outputs against safety, helpfulness, and factual accuracy rubrics and escalated policy violations and edge cases through structured workflows. • Labeled multi-batch dataset items while maintaining high inter-annotator agreement • Cleaned and validated structured datasets (deduplication, schema and formatting verification) before handoff to ML engineering • Built accuracy and throughput reports (e.g., PivotTables/VLOOKUP) to track annotation performance across batches • Evaluated alignment outputs using safety and quality criteria for iterative improvements

2023 - Present

LLM Response Evaluation & Alignment (Selected Project)

OtherTextText

Evaluated AI outputs for helpfulness, harmlessness, and instruction-following dimensions for a major language model alignment effort. Documented evaluation findings to feed directly into alignment iteration cycles and improve model behavior. Built Excel and Google Sheets trackers to monitor annotation accuracy and quality trends across batches. • Assessed model responses against helpfulness/harmlessness/instruction-following criteria • Fed evaluation results into ongoing alignment iteration cycles • Tracked annotation accuracy, batch completion rates, and quality trends over time • Used structured reporting to support continuous improvement of evaluation standards

2024 - 2024

Text Classification & Sentiment Analysis (Selected Project)

OtherTextText

Labeled large text datasets across sentiment and category dimensions using consistent rubric-based judgment to reduce training-set noise. Monitored and reported label distributions and inter-annotator agreement trends to inform project leads and dataset quality. Produced structured trackers to support accuracy and batch completion monitoring over time. • Applied standardized rubric criteria to label sentiment/category • Generated summary charts in Excel and Google Sheets for distribution and agreement trends • Supported quality monitoring for annotation batches via trackers • Contributed dataset cleanliness improvements prior to downstream use

2023 - 2023

Education

U

University of Florida

Bachelor of Science, Computer Science

Bachelor of Science
2018 - 2022
C

Coursera

Certificate in Data Literacy and SQL, Data Literacy and SQL

Certificate in Data Literacy and SQL
Not specified

Work History

V

Various Platforms

UI/UX Designer and QA Contributor

N/A
2024 - Present
V

Various Platforms

AI Evaluation and Data Quality Contributor

N/A
2023 - Present