For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
W
William F.

William F.

AI Evaluator & Red Teamer, Handshake AI (Remote)

USA flagRemote, Usa

Key Skills

Software

Don't disclose
MercorMercor

Top Subject Matter

Vision-Language Models
LLM evaluation
adversarial prompt QA

Top Data Types

TextText
ImageImage

Top Task Types

Red TeamingRed Teaming
RLHFRLHF

Freelancer Overview

AI Evaluator & Red Teamer, Handshake AI (Remote). Brings 3+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Don't disclose and Mercor. Education includes Master of Science, Rensselaer Polytechnic Institute (2027) and Bachelor of Science, Rensselaer Polytechnic Institute (2026). AI-training focus includes data types such as Text and Image and labeling workflows including Red Teaming, Evaluation, and Rating.

Labeling Experience

Mercor

Generalist, Mercor Intelligence (Remote)

MercorMercorTextText

Applied comprehensive aesthetic and structural rubrics to evaluate generative model outputs, focusing on visual design, typographic formatting, spatial layout, and content coherence. Used these evaluations to help improve a top AI lab’s generative models. Ensured labeled evaluations supported consistent assessment of output quality and integrity. • Assessed aesthetic quality and structural integrity using defined rubrics • Evaluated generative model responses for coherence and formatting • Provided quality feedback to improve generation performance • Supported iterative model improvement based on rubric scoring

2026 - Present

AI Evaluator and Red Teamer - Handshake AI

ImageImageRLHFRLHFRed TeamingRed Teaming

Worked as an AI evaluator and red teamer performing adversarial testing on vision-language models. Designed and executed high-complexity visual reasoning prompts to uncover boundary logic failures. Conducted prompt quality assurance and RLHF evaluation to assess response reliability and algorithmic correctness. • Executed adversarial red teaming using dense STEM diagrams and data visualizations • Performed QA on 200+ adversarial tasks to identify edge cases and ambiguity • Evaluated RLHF outputs for logic and Big-O time/space complexity • Tested model robustness against external-domain knowledge bias and inconsistencies.

2025 - Present

AI Evaluator & Red Teamer, Handshake AI (Remote)

Don't discloseTextTextRed TeamingRed Teaming

Performed adversarial red teaming on Vision Language Models using complex visual reasoning prompts with dense STEM diagrams and data visualizations to find boundary logic failures. Conducted prompt QA on over 200 human-authored adversarial tasks to identify edge cases and ensure visual reasoning tests were unambiguous, self-contained, and unbiased by external domain knowledge. Evaluated LLM responses under RLHF by checking logic and Big-O time/space complexity across algorithmic domains (e.g., dynamic programming and greedy algorithms). • Built and reviewed adversarial prompt sets for robustness testing • Audited responses for correctness against expected reasoning criteria • Assessed ambiguity and dependency on external knowledge • Documented findings from evaluation and red teaming cycles

2025 - Present

Education

R

Rensselaer Polytechnic Institute

Master of Science, Information Technology

Master of Science
2026 - 2027
R

Rensselaer Polytechnic Institute

Bachelor of Science, Computer and Systems Engineering

Bachelor of Science
2022 - 2026

Work History

M

Mercor Intelligence

Generalist

Remote
2026 - Present
H

Handshake AI

AI Evaluator and Red Teamer

Remote
2025 - Present