For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
林杨

林杨

AI Data Annotator – NLP, Text Classification & LLM Evaluation

Hong Kong flagN/A, Hong Kong

Key Skills

Software

DoccanoDoccano
Label StudioLabel Studio
ProdigyProdigy
Scale AIScale AI
ArgillaArgilla

Top Subject Matter

NLP Text Annotation and Semantic Classification
LLM Response Evaluation and AI Training Data
Bilingual Content Review and Cross-cultural Communication

Top Data Types

TextText
ImageImage
DocumentDocument

Top Task Types

ClassificationClassification
Data CollectionData Collection
RLHFRLHF
SegmentationSegmentation
PolygonPolygon
Evaluation/RatingEvaluation/Rating
Computer Programming/CodingComputer Programming/Coding

Freelancer Overview

I have practical experience in AI training data preparation, text annotation, and linguistic quality review through my CBS5502 NLP project. I worked with the Kaggle MBTI English corpus, where we processed 8,675 raw records into over 421,000 post segments and extracted target-word sentences for “think” and “feel.” For the “feel” subset, we identified 22,587 candidate sentences, retained 18,636 usable items, and built a manually reviewed gold set for semantic classification. My strongest fit is text-based AI training work, including classification, evaluation, data cleaning, and guideline-based review. I reviewed “feel” sentences by distinguishing experiential/state meanings from propositional/judgment meanings, confirming labels, correcting errors, excluding noisy or incomplete examples, and documenting ambiguous cases. With my background in computational linguistics, generative AI, bilingual communication, and design QA, I am detail-oriented, consistent with annotation guidelines, and able to handle nuanced language data for LLM evaluation and NLP training tasks.

Labeling Experience

Disambiguating “Feel” in MBTI Thinking vs. Feeling Types

TextTextEvaluation/RatingEvaluation/Rating

In my CBS5502 final project, I worked on an NLP data preparation and manual annotation workflow using the Kaggle MBTI English corpus. The project processed 8,675 raw rows into over 421,000 post segments, extracted target-word sentences for “think” and “feel,” and produced cleaned datasets with traceable audit logs. For the “feel” subset, we identified 22,587 candidate sentences, kept 18,636 usable items, and created a balanced 400-row gold set for manual review. My annotation work focused on distinguishing A: experiential/state uses from B: propositional/judgment uses of “feel.” I followed detailed review guidelines, checked whether each sentence should remain in the gold set, confirmed or corrected labels, marked exclusions such as truncated fragments or non-target usage, and supported final quality control. This experience strengthened my ability to follow annotation standards, resolve linguistic ambiguity, document decisions, and produce reliable training data for AI/NLP systems.

2026 - 2026

Education

H

Hong Kong polytechnic university

master, Artificial Intelligent and Humanity

master
2025 - 2026

Work History

U

U-FEI Environmental Technology Co., Ltd.

Designer / Project Manager

dongguan
2023 - 2025