For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
D
Delaine J.

Delaine J.

Expert AI Evaluator & RLHF Specialist | Outlier by Scale AI (Remote)

USA flagSan Francisco, Usa

Key Skills

Software

Scale AIScale AI
Data Annotation TechData Annotation Tech
Other
TolokaToloka

Top Subject Matter

LLM evaluation and RLHF training data for coding/reasoning
AI safety alignment
Red-teaming Domain Expertise

Top Data Types

TextText
AudioAudio

Top Task Types

RLHFRLHF
Red TeamingRed Teaming
Fine-tuningFine-tuning

Freelancer Overview

Expert AI Evaluator & RLHF Specialist | Outlier by Scale AI (Remote). Brings 4+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Scale AI, Data Annotation Tech, and Other. Education includes Master of Science, Stanford University (2025) and Bachelor of Science, University of California, Los Angeles (UCLA) (2023). AI-training focus includes data types such as Text and labeling workflows including RLHF, Red Teaming, and Evaluation.

Labeling Experience

Data Annotation Tech

AI Safety & Red-Teaming Specialist | DataAnnotation.Tech (Remote)

Data Annotation TechData Annotation TechTextTextRed TeamingRed Teaming

Conducted adversarial testing to generate and validate jailbreak prompts that attempt to bypass model safety filters and operational guidelines. Documented vulnerability vectors and provided actionable feedback to patch alignment flaws. Performed exhaustive fact-checking on technical model outputs to minimize hallucinations. • Built sophisticated jailbreak prompts for safety filter circumvention testing • Recorded vulnerability patterns such as roleplay circumvention, payload splitting, and base64 encoding attacks • Fact-checked outputs using robust research methodologies in technical domains • Supplied engineering-ready guidance to remediate alignment vulnerabilities

2024 - Present
Scale AI

Expert AI Evaluator & RLHF Specialist | Outlier by Scale AI (Remote)

Scale AIScale AITextTextRLHFRLHF

Performed RLHF-oriented evaluation of LLM-generated code and reasoning for training-data generation. Created multi-turn conversational prompt datasets to test memory, contextual awareness, and instruction-following over extended contexts. Produced detailed critiques and rankings to support gradient-descent optimization and model refinement. • Evaluated and ranked LLM outputs containing Python and C++ code plus logic chains • Wrote and executed multi-turn prompts for long-context behavioral testing • Delivered justification-style feedback for response ranking and iterative refinement • Generated high-fidelity domain-expert training signals for foundational models

2024 - Present

LLM Benchmarking Analyst | Handshake AI (Remote)

OtherTextText

Benchmarked experimental LLMs by running automated and manual test suites for quantitative reasoning and coding accuracy. Analyzed performance data to identify model weaknesses including transformer attention and context-window degradation. Drafted evaluation reports describing regressions between deployment versions. • Executed rigorous automated/manual benchmark suites for coding and NLP capability • Interpreted metrics to locate architectural weaknesses in attention and long-context behavior • Produced structured evaluation reports for model comparison across versions • Identified edge-case failure modes impacting reasoning and coding quality

2023 - 2024
Toloka

Data Annotator & QA Lead | Remotasks & Toloka (Remote)

TolokaTolokaTextTextFine-tuningFine-tuning

Generated and labeled complex domain-specific datasets used for supervised fine-tuning of enterprise AI models. Performed advanced QA on junior annotator outputs, enforcing adherence to labeling rubrics and maintaining high acceptance quality. Categorized edge-case responses and mapped logical fallacies and syntax errors to support reward-model refinement for RLHF training. • Labeled training data for SFT pipelines in next-generation enterprise models • Conducted QA reviews against project rubrics to ensure consistency and accuracy • Flagged and categorized edge cases including logical fallacies and syntax errors • Supported engineering iteration by providing structured feedback for reward model tuning

2023 - 2023

Education

S

Stanford University

Master of Science, Computer Science

Master of Science
2023 - 2025
U

University of California, Los Angeles (UCLA)

Bachelor of Science, Computer Science

Bachelor of Science
2020 - 2023

Work History

T

TechNova Solutions

Software Engineer, AI Systems

San Francisco
2025 - Present
S

Stanford AI Lab

Graduate Research Assistant, AI

Stanford
2023 - 2025