For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
I
Idris B.

Idris B.

AI Quality & Evaluation Specialist

USA flagRock Hill, Usa

Key Skills

Software

No software listed

Top Subject Matter

Operations & Process Improvement
Model Training & Evaluation
Rubric Design

Top Data Types

ImageImage
TextText
VideoVideo

Top Task Types

Data CollectionData Collection
Red TeamingRed Teaming
RLHFRLHF
Fine-tuningFine-tuning
Evaluation/RatingEvaluation/Rating
MappingMapping

Freelancer Overview

IDRIS E’JAIZ BAILEY Generative AI Safety · Model Evaluation & Red Teaming · Strategy & Operations [email protected] · +1 704 231 7662 · Eastern Time Zone · Remote · US Citizen · 100% T&P Veteran PROFILE Trained auditor, operator, and writer who treats model evaluation as frontline safety work - every failure caught upstream is one a user never has to encounter. Across 70+ evaluations of generative AI output, I write, review, and rank model responses against detailed rubrics and safety-related categories with one common goal in mind: identifying where a large language model turns unsafe, noncompliant, inconsistent, or otherwise problematic, isolating the point of failure, and aligning each to its defect category. I construct and probe varied test prompts, hold sound judgment steady on ambiguous and policy-sensitive outputs through high-volume sprints, and write the structured rationale that feeds model-training calibration. The bigger picture is the part I care about : frontier models reach millions of people, and the line between helpful and harmful is set upstream by the quality and consistency of human judgment applied to model behavior - the exact discipline eight years of military and operational auditing built into me. Now concentrating that discipline on AI safety and red-team evaluation. Domain depth spans finance, military operations, cross-border legal and regulatory, and consumer business. PROFESSIONAL EXPERIENCE AI Trainer & Data Annotator · M

Labeling Experience

Multi-Rubric Artifact Evaluation & Quality Assurance

DocumentDocumentEvaluation/RatingEvaluation/Rating

Completed 140++ artifact evaluations at production sprint pace across three distinct rubric types: aesthetic ranking, style-matching, and strict 1:1 template-matching criteria. Deliver structured multi-paragraph written rationales documenting specific observations (clipping, alignment failures, hierarchy breaks, spacing inconsistencies) with reasoning calibrated to rubric criteria. Achieve zero revisions on submitted work by maintaining rubric consistency across high-volume output, controlling for relative-comparison bias in head-to-head rankings, and applying criteria independently per response. Operate within defined time targets (three full artifact reviews per two-hour working block, ~5 minutes per evaluation against 8-minute ceiling). Identify edge cases and ambiguous criteria; surface patterns warranting rubric refinement to support downstream inter-rater calibration. Experience on 23,000-person evaluation project.

2026 - Present

Browser Agent Benchmark Design & Evaluation Authorship

TextTextPrompt + Response Writing (SFT)Prompt + Response Writing (SFT)

Authoring outcome-focused benchmark tasks for evaluating AI browser agent capability on web-based research and interaction workflows. Supporting 500-task benchmark spanning three capability categories: targeted lookup, form/filter interaction, and hybrid research+interaction tasks. Each task includes: concrete user-intent prompt, task metadata (primary website, starting URL, capability category, output format), success criteria (atomic, binary, outcome-measured), and grading rubric. Executed each task independently according golden response template; validated alignment between prompt, success criteria, and rubric through automated QC gates. Emphasis: outcome-focused evaluation (what the agent accomplished) over path-focused grading (how the agent solved it). Quality discipline: one constraint per criterion, every criterion checks full result set, no hard-coded values, live-web-proof validation.

2026 - Present

Education

H

HEC Paris and a bachelor's in Economics from Howard University with a minor in Business Administration.

I have an MBA in Finance

I have an MBA in Finance
Not specified

Work History

C

Company not specified

Operations & Quality Analysis | finmid GmbH Conducted process audits and quality control analysis across operational wor

Location not specified
Not specified