For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
胡文祥

胡文祥

NER/Punctuation Restoration Data Labeling Specialist

China flagWuHan, China

Key Skills

Software

Label StudioLabel Studio

Top Subject Matter

Classical Chinese text segmentation and punctuation restoration
Banking document archival and retrieval
HR document information extraction

Top Data Types

TextText
DocumentDocument
ImageImage

Top Task Types

Entity (NER) ClassificationEntity (NER) Classification
ClassificationClassification
Bounding BoxBounding Box
SegmentationSegmentation
RLHFRLHF
Text GenerationText Generation
Question AnsweringQuestion Answering
Evaluation/RatingEvaluation/Rating

Freelancer Overview

Classical Chinese Punctuation Restoration NER Data Labeling. Brings 11+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Internal and Proprietary Tooling. Education includes Bachelor of Engineering, Beijing Jiaotong University (2021). AI-training focus includes data types such as Text and Document and labeling workflows including Entity (NER) Classification and Classification.

Labeling Experience

Government Organization Event Classification Labeler

DocumentDocumentClassificationClassification

The project consisted of annotating and classifying government documents for organization-change event detection using LLM systems. Annotators labeled event categories for institution establishment, restructuring, renaming, and cancellation. The labeled dataset enabled accurate organizational change recognition and timeline construction for event visualization. • Annotated organization-change events across diverse government documents. • Provided classification labels for fine-tuning the LLM extraction pipeline. • Contributed labeled data for organization-centered timeline modeling. • Utilized internal/proprietary annotation technology during labeling.

2024 - Present

Personnel Archive Event Classification Labeler

DocumentDocumentClassificationClassification

Classified HR documents into event types such as appointment, promotion, demotion, resignation, and dismissal using large language models. Labeled document data for training and evaluation of LLM-based event extraction systems. Facilitated the construction of person-centered knowledge graphs powered by labeled personnel event data. • Applied LLM models to enterprise personnel archives. • Developed labeled datasets for event extraction tasks. • Supported knowledge graph building with labeled data. • Focused on HR events and related changes.

2024 - Present

HR Document Field Extraction Data Labeler

DocumentDocumentEntity (NER) ClassificationEntity (NER) Classification

The experience focused on field-level labeling for employee document digitization using multimodal document understanding. Annotation responsibilities included marking over 100 types of HR-related fields to support accurate model-driven extraction and structured data archiving. Labeled data enabled the introduction of human review processes before final database ingestion. • Annotated personnel field data from diverse HR archive documents. • Prepared training data for multimodal document understanding models. • Facilitated downstream human review workflows with comprehensive annotations. • Leveraged proprietary/enterprise-level annotation software.

2022 - 2022

Banking Document Alignment Labeling Specialist

DocumentDocumentClassificationClassificationEntity (NER) ClassificationEntity (NER) Classification

This project centered on extracting and aligning key fields from scanned paper documents and their digital equivalents in the banking domain. Human reviewers annotated document entities such as IDs, account numbers, and amounts for model training and archival support. Annotated data were used for field extraction, database matching, and downstream pipeline integration. • Labeled document entities for both paper and digital banking records. • Supported OCR pipeline development through ground truth annotation. • Annotated structured fields to facilitate accurate digital archiving. • Utilized internal annotation systems and validation tools.

2020 - 2020

NER/Punctuation Restoration Data Labeling Specialist

TextTextEntity (NER) ClassificationEntity (NER) Classification

This labeling project involved identifying sentence boundaries and restoring punctuation in classical Chinese texts using a BERT-based NER approach. The task improved the readability of historical documents and enabled downstream information extraction. The process required annotation of segmentation points and punctuation types for training and evaluation of the model. • Labeled unpunctuated classical Chinese texts for sentence boundary detection. • Assigned punctuation type labels to text spans for machine learning training. • Validated labeled data to support NER model development. • Used proprietary/internal annotation tooling for efficient workflow.

2019 - 2019

Education

B

Beijing Jiaotong University

Bachelor of Engineering, Engineering Management

Bachelor of Engineering
2021 - 2021

Work History

H

Hanwang Data Technology

AI R&D Engineer

Wuhan
2016 - Present