For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
G
Gm W.

Gm W.

AI Response Evaluator / Model Response Quality Reviewer (LLM evaluation)

China flagN/A, China

Key Skills

Software

Other

Top Subject Matter

LLM/model response evaluation
Chinese content QA
user-intent grounded assessment

Top Data Types

TextText

Top Task Types

Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
RLHFRLHF

Freelancer Overview

AI Response Evaluator / Model Response Quality Reviewer (LLM evaluation). Brings 5+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Claude, Claude Code, and Other. Education includes Bachelor of Engineering, Jiangsu Ocean University (2016). AI-training focus includes data types such as Text and labeling workflows including Evaluation, Rating, and Prompt + Response Writing (SFT).

Labeling Experience

AI Data Annotator / Preference Ranking & Prompt/Response Evaluation

OtherTextTextRLHFRLHF

Conducted preference-ranking style evaluation of multiple AI responses and tested prompts to determine which outputs better satisfy user requirements. Applied strong Chinese language understanding to assess content categorization, harmful/low-quality detection, sentiment or stance judgment, and summary quality. Produced guideline-aligned labeling rationales and error tagging to support high-quality data creation for model improvement. • Rated and compared responses for usefulness, safety, and alignment with intent • Classified user intent and content categories; assessed sentiment/stance and summary quality • Detected harmful or low-quality content and documented error rationales • Tagged errors and recommended prompt or response changes for iterative training

2022 - 2024

AI Trainer / AI-Assisted Coding Evaluator (Claude Code workflow)

TextTextPrompt + Response Writing (SFT)Prompt + Response Writing (SFT)

Performed prompt refinement and iterative task decomposition to guide AI-assisted coding and model behavior improvements. Reviewed AI-generated code changes by checking diffs, verifying behavior, and flagging logic gaps or over-implementation issues. Provided coding-related evaluation feedback using structured review comments that translate requirements into actionable fixes. • Guided AI with natural-language requirements and prompt refinement for coding tasks • Reviewed AI code edits, verified behavior, and evaluated bug-fix explanations • Detected common AI failure modes such as misunderstanding requirements and hallucinated APIs • Delivered intermediate-level assessment of technical task plans and generated outputs

2017 - 2021

AI Response Evaluator / Model Response Quality Reviewer (LLM evaluation)

TextText

Provided AI response evaluation and quality feedback for LLM outputs, focusing on factual accuracy, instruction-following, completeness, reasoning consistency, and safety risk identification. Used product and user-intent analysis skills to judge whether outputs truly solve the user need and to rank improvements based on usefulness. Produced structured, evidence-based review comments and error classifications to support iterative refinement of model responses. • Evaluated whether responses were accurate, complete, useful, safe, and aligned with user intent • Identified missing steps, unclear logic, misleading claims, and hallucinations • Wrote actionable improvement suggestions and rationale-style feedback • Supported preference ranking and RLHF-style evaluation workflows

2012 - 2016

Education

J

Jiangsu Ocean University

Bachelor of Engineering, Communication Engineering

Bachelor of Engineering
2012 - 2016

Work History

P

Publicly Listed Internet Company

Product Manager

N/A
2019 - 2023