For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
激推

激推

Text Classification Reward Model with Qwen2-3B and Chain-of-Thought Signals LLM / Reward Modeling

HongKong, A

Key Skills

Software

AWS SageMakerAWS SageMaker
Anno-MageAnno-Mage

Top Subject Matter

LLM evaluation and reward modeling for reasoning-data quality
Academic text review and publication quality control
Academic peer review and LLM-assisted review generation

Top Data Types

TextText
DocumentDocument

Top Task Types

Fine-tuningFine-tuning
Text GenerationText Generation
Data CollectionData Collection

Freelancer Overview

Text Classification Reward Model with Qwen2-3B and Chain-of-Thought Signals LLM / Reward Modeling. Brings 5+ years of professional experience across legal operations, contract review, compliance, and structured analysis. Core strengths include PyTorch, Transformers, and LoRA. Education includes Master of Engineering Management, Peking University (2026) and Bachelor of Science, Northwest A&F University (2020). AI-training focus includes data types such as Text and Document and labeling workflows including Fine-tuning, Evaluation, and Rating.

Labeling Experience

MDPI Beijing | SCIE Journal Section Managing Editor

DocumentDocument

Managed an academic journal section workflow that required organizing and processing manuscript submissions into structured publication decisions. Oversaw topic-focused special-issue projects and coordinated review pipeline operations to support editorial quality control. Maintained large-scale scholar and board networks to ensure consistent manuscript-source handling and evaluation. • Managed SCIE-indexed journal section operations and annual business plans • Developed and managed 40+ frontier-topic special issues for publication • Coordinated review pipeline at scale, publishing 200+ papers • Maintained 1,600+ editorial-board and scholar network for quality assurance

2021 - 2024

Text Classification Reward Model with Qwen2-3B and Chain-of-Thought Signals LLM / Reward Modeling

TextTextFine-tuningFine-tuning

Built and evaluated an LLM reward-modeling workflow to classify question quality and verify reasoning logic. The work supported generation of higher-quality reasoning training data via reward modeling. Implemented fine-tuning and evaluation experiments to measure classification performance and improve reasoning-data quality. • Reward model workflow on Qwen2-3B for reasoning logic and question-quality evaluation • Fine-tuned with LoRA and 4-bit quantization using PyTorch and Transformers • Evaluated with F1 and precision metrics to validate label quality • Aligned outputs to LLM evaluation, reasoning step verification, and preference/reward data pipelines

2024

CogniTrace | Proof-of-Process Infrastructure for Human-Led Writing

TextTextData CollectionData Collection

Built a proof-of-process capture and scoring system that collects local writing-process signals and exports evidence for explainable human-led production. The extension analyzes session behavior and produces timeline replay and integrity scoring outputs. This enables process-evidence labeling to support quality control and reduce limitations of output-only AI detection. • Developed Chrome extension MVP for Google Docs and Word Online process capture • Designed timeline replay and integrity scoring to quantify process evidence • Implemented MVP certificate export for auditable writing provenance • Targeted explainable evidence workflows for human-led production verification

2021

ScholarReview-LLM | LLM-Based Academic Paper Review System

TextTextText GenerationText Generation

Developed an LLM-based academic paper review system to generate structured review outputs from submitted documents. The system processes PDFs/DOCX/Markdown, runs model inference, and produces organized review content for human editing. The work supports annotation and quality-control tasks for academic text by standardizing review generation. • Built Gradio paper-review application and single-paper review scripts • Implemented PDF/DOCX/Markdown processing and structured review generation • Organized thesis materials, datasets, model outputs, and code for reproducibility • Refactored repository structure to improve environment-based hygiene and consistency

2021

Education

P

Peking University

Master of Engineering Management, Master of Engineering Management

Master of Engineering Management
2024 - 2026
U

University of Nebraska-Lincoln

Bachelor of Science, Food Science and Technology

Bachelor of Science
2019 - 2020

Work History

L

Lakefront Asset Management

Investor Relations Intern

Beijing
2024 - 2025
C

China Merchants Securities

Analyst Intern

Beijing
2024 - 2024