For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
S
Shawn W.

Shawn W.

Artificial Intelligence Engineer — Multimodal translation dataset labeling/training support at Yunnan Provincial Key Lab

Singapore flagSingapore, Singapore

Key Skills

Software

No software listed

Top Subject Matter

Multimodal machine translation (NLP)
AI video auto-clip metadata generation/training data preparation
Content moderation / safety model training

Top Data Types

DocumentDocument
VideoVideo
TextText
ImageImage

Top Task Types

Function CallingFunction Calling
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
Fine-tuningFine-tuning

Freelancer Overview

Artificial Intelligence Engineer — Multimodal translation dataset labeling/training support at Yunnan Provincial Key Lab. Brings 6+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Internal, Proprietary Tooling, and Docker. Education includes Bachelor of Science, Guilin University of Electronic Technology (2018) and Master of Science, Kunming University of Science and Technology (2024). AI-training focus includes data types such as Document, Video, and Text and labeling workflows including Translation, Localization, and Functi

Labeling Experience

Artificial Intelligence Engineer — LLM fine-tuning on government documents at Kunming Information Hub Co.

DocumentDocumentFine-tuningFine-tuning

Fine-tuned a large language model on government document samples using instruction-tuning and preference/optimization methods. The effort distilled 10K+ state council report samples into a training dataset and deployed the resulting model for inference. Training used SFT plus GRPO and was served via vLLM for practical usage.• Distilled 10K+ government report samples for model training.• Trained Qwen3.5-9B using SFT + GRPO for improved task behavior.• Prepared structured text inputs/outputs for instruction tuning.• Deployed the fine-tuned model via vLLM for downstream document interactions.

2024 - Present

Artificial Intelligence Engineer — Content-safety model fine-tuning and dataset refactoring at Kunming Information Hub Co.

TextTextPrompt + Response Writing (SFT)Prompt + Response Writing (SFT)

Rebuilt a content-safety training workflow by refactoring a labeled sample set and iterating multiple safety model versions. The process involved preparing 48K safety samples and running supervised fine-tuning and related training/evaluation cycles to meet safety benchmarks. Results reached 80%+ F1 and the system protected 1M+ users in production.• Refactored 48K safety samples for training and evaluation.• Iterated through 4 model versions to improve benchmark performance.• Achieved 80%+ F1 across content-safety benchmarks.• Supported safe deployment targeting protection for 1M+ users.

2024 - Present

Artificial Intelligence Engineer — AI video auto-clip pipeline (metadata extraction, indexing, dedup) at Kunming Information Hub Co.

VideoVideoFunction CallingFunction Calling

Delivered an end-to-end AI video auto-clip pipeline by producing structured VLM-derived metadata for clip selection. This required organizing and generating consistent metadata fields for indexing and retrieval. The system performed content deduplication and production deployment for scalable processing.• Generated VLM metadata from video segments to drive auto-clip behavior.• Built SHA256-based deduplication for consistent dataset indexing/metadata management.• Stored and searched embeddings using MinIO + OpenSearch vector DB integration.• Deployed the pipeline in Docker for production-grade automated processing.

2024 - Present

Artificial Intelligence Engineer — Multimodal translation dataset labeling/training support at Yunnan Provincial Key Laboratory of Artificial Intelligence

DocumentDocument

Built a multimodal translation pipeline using paired sentence data to improve translation quality and model adequacy. Labeled/paired text instances for training and evaluation using multi-sentence corpora. The work included dataset curation and measurement against BLEU and human adequacy ratings.• Multimodal translation training with mCLIP + mBART using 50K+ sentence pairs.• Improved translation metrics (+10 BLEU, +2.4 adequacy rating) through iterative data/pipeline refinement.• Prepared paired input-output text records for downstream evaluation and deployment.• Used translation task outputs as structured targets for model optimization.

2021 - 2024

Education

K

Kunming University of Science and Technology

Master of Science, Machine Learning

Master of Science
2021 - 2024
G

Guilin University of Electronic Technology

Bachelor of Science, Computer Science

Bachelor of Science
2014 - 2018

Work History

K

Kunming Information Hub Co., Ltd.

Artificial Intelligence Engineer

Kunming
2024 - Present
Y

Yunnan Provincial Key Laboratory of Artificial Intelligence

Artificial Intelligence Engineer

Kunming
2021 - 2024