For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
S
Sean B.

Sean B.

AI Model Evaluator & Annotator – LLM Alignment, Red Teaming & Code Evaluation

USA flagSacramento, Usa

Key Skills

Software

AppenAppen
ClickworkerClickworker
CloudFactoryCloudFactory
CVATCVAT
CrowdSourceCrowdSource
Data Annotation TechData Annotation Tech
Google Cloud Vertex AIGoogle Cloud Vertex AI
Figure EightFigure Eight
HiveMindHiveMind
HumanaticHumanatic
iMeritiMerit
Img Lab
LabelboxLabelbox
MercorMercor
Micro1
MindriftMindrift
OneFormaOneForma
OpenCV AI Kit (OAK)OpenCV AI Kit (OAK)
RemotasksRemotasks
SuperAnnotateSuperAnnotate
Snorkel AISnorkel AI
Scale AIScale AI
Surge AISurge AI
TolokaToloka
V7 LabsV7 Labs
Don't disclose
Internal/Proprietary Tooling
RoboflowRoboflow
SuperviselySupervisely

Top Subject Matter

Artificial Intelligence & Machine Learning – LLM Evaluation & Alignment
Software Engineering & Computer Science – Code Annotation & Prompt Engineering
Cybersecurity & AI Safety – Adversarial Red-Teaming & Failure Mode Detection

Top Data Types

TextText
Computer Code ProgrammingComputer Code Programming
ImageImage

Top Task Types

Bounding BoxBounding Box
PolygonPolygon
ClassificationClassification
Entity (NER) ClassificationEntity (NER) Classification
Point/Key PointPoint/Key Point
Object DetectionObject Detection
Text GenerationText Generation
Question AnsweringQuestion Answering
Text SummarizationText Summarization
RLHFRLHF
Fine-tuningFine-tuning
Red TeamingRed Teaming
TranscriptionTranscription
Evaluation/RatingEvaluation/Rating
Computer Programming/CodingComputer Programming/Coding
Data CollectionData Collection
Function CallingFunction Calling
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
SegmentationSegmentation

Freelancer Overview

Sean Boggs AI Specialist & Software Engineer 530-417-3999 | [email protected] | Rancho Cordova, CA | github.com/DiggityDooo PROFESSIONAL OVERVIEW Results-driven AI Specialist and full-stack software engineer with hands-on experience in model evaluation, RLHF annotation, adversarial red-teaming, and instruction tuning across multiple AI platforms. Combines a cognitive psychology academic background with deep technical proficiency in Python, TypeScript, and cloud-native architectures to deliver precise, multi-dimensional AI feedback at scale. Proven ability to identify model failure modes, engineer targeted alignment strategies, and ship production software tools — from open-source OSINT CLIs to Claude API-powered web applications. WORK EXPERIENCE AI TRAINER & AI SPECIALIST | AfterQuery, RemoteMAY 2026 – PRESENT • Spearheaded generative AI model evaluation and instruction tuning, applying cognitive science principles to systematically debug model hallucinations and implement targeted alignment strategies. • Designed and executed adversarial prompts to stress-test model behavior across reasoning, coding, and instruction-following domains, delivering structured feedback that improved output quality and reduced failure rates. • Collaborated cross-functionally with research teams to develop alignment strategies, leveraging full-stack engineering expertise to provide technically rigorous, reproducible model assessments. AI MODEL EVALUATOR | Mercor AI, Remote (Contract)JAN 2026

Labeling Experience

Code Prompt Engineering & Instruction Tuning — AI Training Data Production

Computer Code ProgrammingComputer Code ProgrammingComputer Programming/CodingComputer Programming/Coding

Engineered high-quality prompt-response pairs and instruction-tuning examples for code generation model training, applying full-stack software engineering expertise across Python, TypeScript, JavaScript, and Bash. Designed precise, unambiguous task prompts covering algorithm implementation, API usage, debugging, and system design — and produced corresponding gold-standard solutions with structured annotations covering correctness, efficiency, style adherence, and edge-case handling. Contributed to SFT and RLHF pipelines by generating reproducible adversarial code prompts designed to surface model failure modes in reasoning-heavy and multi-step implementation tasks.

2026 - Present

LLM Adversarial Red-Teaming — Hallucination Detection & Safety Evaluation

TextTextRed TeamingRed Teaming

Designed and executed systematic adversarial prompt suites to stress-test large language model behavior across reasoning, coding, factuality, and safety domains. Identified and documented latent edge-case failure modes including hallucinations, instruction-following breakdowns, jailbreak vulnerabilities, and unsafe output patterns. Delivered structured failure reports with reproducible test cases used by research teams to target alignment interventions. Maintained a rigorous red-team methodology grounded in cognitive behavioral analysis, systematically probing model responses for bias, sycophancy, and policy violations at production scale.

2026 - Present

RLHF Preference Labeling — LLM Response Ranking & Pairwise Comparison

TextTextRLHFRLHF

Executed production-scale RLHF annotation including pairwise response comparison, ranked preference labeling, and multi-criteria quality scoring across LLM outputs spanning reasoning, instruction-following, mathematics, and dialogue tasks. Applied fine-grained rubrics evaluating helpfulness, harmlessness, honesty, groundedness, and stylistic adherence. Generated high-quality gold-standard training signals that fed directly into model alignment and performance optimization pipelines. Maintained high annotation throughput with strict inter-rater agreement standards and minimal supervision across ambiguous, open-ended labeling tasks at production scale.

2026 - Present

Multimodal Image Quality Evaluation — AI Model Training

ImageImageEvaluation/RatingEvaluation/Rating

Performed high-volume image quality evaluation and relevance scoring for multimodal AI model training pipelines. Tasks included visual accuracy assessment, bounding-region relevance scoring, image-response alignment rating, and multi-dimensional quality rubric scoring across helpfulness, visual groundedness, and safety compliance. Produced structured calibration datasets and curated golden-answer sets that directly improved inter-rater agreement scores and downstream model alignment quality. Maintained strict guideline adherence and high annotation throughput across diverse visual task types throughout a AI Fellowship program.

2025 - Present

Education

S

Sacramento City College

Concurrent Enrollment alongside FLC, Psychology

Concurrent Enrollment alongside FLC
2024 - 2026
F

Folsom Lake College

Transfer Degree, Psychology/Mechanical Engineering

Transfer Degree
2024 - 2026

Work History

A

AfterQuery

AI data Instructor

Sacramento
2026 - Present
M

Mercor

Artificial Intelligence Researcher/ RedTeamer

Sacramento
2026 - Present