For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
J
James M.

James M.

Data Annotator / LLM Response Evaluator (Independent Contractor, Remote)

USA flagBogalusa, Usa

Key Skills

Software

Data Annotation TechData Annotation Tech
LabelboxLabelbox
OneFormaOneForma

Top Subject Matter

LLM response evaluation
RLHF dataset/rubric creation
and hallucination detection

Top Data Types

TextText
DocumentDocument
ImageImage

Top Task Types

RLHFRLHF
Evaluation/RatingEvaluation/Rating
ClassificationClassification
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
Text GenerationText Generation
Question AnsweringQuestion Answering
Red TeamingRed Teaming
Text SummarizationText Summarization

Freelancer Overview

Expert Data Annotator and LLM Response Evaluator specializing in Reinforcement Learning from Human Feedback (RLHF), adversarial red teaming, and Supervised Fine-Tuning (SFT). Proven track record of testing pre-release, development-stage AI models by engineering complex 8+ turn conversational scenarios and conducting rigorous A/B preference evaluations. Highly skilled in crafting detailed, objective rationales to identify edge cases, behavioral drift, and safety guardrail failures, ensuring models maintain factual grounding and contextual memory across extended interactions. Brings over 15 years of professional experience, distinguishing myself through specialized domain expertise in financial risk assessment, regulatory compliance, and tax frameworks. This is supported by former Series 6 and insurance licenses, alongside an Advanced Volunteer Income Tax Assistance (VITA) certification. This analytical background allows for the execution of highly technical hallucination audits, policy compliance testing, and quantitative logic evaluations that standard annotators cannot perform. Education includes a Master of Science from Regis University (2008) and a Bachelor of Science from the University of Southern Mississippi (2003), combining academic rigor with a meticulous, quality-focused approach to continuous AI model alignment.

Labeling Experience

Data Annotator / LLM Response Evaluator (Independent Contractor, Remote)

OtherTextTextRLHFRLHF

Worked as an LLM data annotator and response evaluator for RLHF-style optimization by scoring outputs against multiple task criteria and safety requirements. Created and calibrated high-fidelity prompt training sets and detailed 20-point technical evaluation rubrics to ensure response quality and logical consistency. Audited model outputs against binary guidelines and used verification techniques to isolate errors such as factual issues and hallucinations. • Scored responses on Instruction Following, Completeness, Relevance, Accuracy, and Safety • Generated prompt training sets and 20-point technical evaluation rubrics for calibration • Audited outputs to identify hallucinations, bias, punts, and refusal-to-answer failures • Performed comparative analysis and provided preference justifications for response ranking

2025 - Present

Adversarial Red Teamer (Finance & Hallucination Audits)

TextTextRed TeamingRed Teaming

Scope: Conducted targeted adversarial red teaming to expose vulnerabilities and behavioral flaws in large language models, with a specific focus on hallucination induction within complex quantitative and financial contexts. Tasks Performed: Engineered sophisticated adversarial prompts to stress-test model boundaries and force the generation of factual inaccuracies or definitively wrong answers. Developed targeted test cases designed to manipulate the model into lying or fabricating information to mask its own limitations or lack of capability. Project Size: Executed extensive vulnerability probing as part of a broader hallucination audit and safety alignment initiative. Quality Measures: Meticulously documented successful prompt injection strategies, evasion tactics, and behavioral drifts. Logged detailed failure modes to provide actionable vulnerability reports for safety mitigation and model fine-tuning.

2025 - 2026

Real-Time Multimodal AI Evaluator

TextTextRLHFRLHF

Scope: Conducted comparative A/B testing on pre-release, development-stage large language models to optimize performance, safety, and human alignment prior to deployment. Tasks Performed: Engineered targeted 8+ turn conversational data based on complex scenarios to test model memory and behavioral drift. Evaluated raw model outputs for logical reasoning, consistency, and contextual retention across these extended interactions. Authored highly detailed, objective rationales explaining model failures, identifying edge cases, and providing critical qualitative feedback for engineer-led fine-tuning loops. Project Size: Participated in large-scare, continuous evaluation cycles as part of an ongoing reinforcement learning initiative. Quality Measures: Adhered to strict, rapidly shifting multi-page project rubrics. Consistently maintained high quality-assurance metrics by rigorously verifying factual grounding and ensuring strict adherence to behavioral guardrails.

2025 - 2026

Education

R

Regis University

Master of Science, Management

Master of Science
2005 - 2008
U

University of Southern Mississippi

Bachelor of Science, Hospitality Management

Bachelor of Science
1998 - 2003

Work History

T

The Lake Allstate Insurance and Financial

Risk Advisor

Metairie
2016 - 2017
P

Prudential Financial

Financial Professional Associate

Metairie
2016 - 2016