For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
M
Muye S.

Muye S.

Biomedical AI & LLM Evaluation Specialist | Medical Data, Proteomics, OCR, and Multimodal Annotation

China flagZHEJIANG, China

Key Skills

Software

Label StudioLabel Studio
CVATCVAT
RoboflowRoboflow
Other

Top Subject Matter

Biomedical AI, Medical Imaging, Proteomics & Cancer Research
LLM Evaluation, RLHF, Prompt Review & AI Response Quality Assessment
Multimodal Data Annotation, OCR Information Extraction & Knowledge Graphs

Top Data Types

DocumentDocument
ImageImage
TextText

Top Task Types

Question AnsweringQuestion Answering
Object DetectionObject Detection
Computer Programming/CodingComputer Programming/Coding
Evaluation/RatingEvaluation/Rating
Fine-tuningFine-tuning

Freelancer Overview

I have hands-on experience in AI training data, data annotation, and model evaluation across text, image, OCR, and medical imaging scenarios. My work has involved reviewing model outputs, comparing responses, identifying hallucinations, improving prompt quality, designing annotation standards, and evaluating whether AI-generated answers are accurate, complete, safe, and useful. I am especially strong in Chinese-language data, bilingual Chinese-English content review, structured information extraction, and high-consistency labeling workflows. What sets me apart is that I am not only a data labeler, but also an AI project builder with technical understanding of machine learning, medical image segmentation, multimodal datasets, and automated data pipelines. I have worked with CT image segmentation projects, OCR-based business card extraction, LLM prompt engineering, AI evaluation criteria, and Python/PyTorch-based model workflows. This allows me to understand both the annotation task itself and the downstream impact on model training quality, making me highly suitable for AI training, RLHF, LLM evaluation, medical data review, and multimodal annotation projects.

Labeling Experience

Factor Generation & AI Training Data for Quantitative Research

DocumentDocumentFine-tuningFine-tuning

Worked on an AI-assisted quantitative factor generation and evaluation workflow focused on alpha signal discovery, factor expression generation, simulation review, and quality filtering. The project involved designing and testing candidate factor expressions, reviewing model-generated factor ideas, identifying invalid or low-quality signals, recording failed gates, and building reusable research memory to improve future factor generation. This experience is highly relevant to AI training data because it required evaluating model outputs, creating feedback signals, classifying failures, standardizing candidate quality, and turning research results into structured training/evaluation data. The workflow covered financial time-series data, market features, ranking operators, neutralization logic, turnover control, drawdown review, Sharpe/Fitness-style metrics, and automated candidate generation pipelines. I also worked with local templates, LLM-assisted generation, JSONL queues, simulation logs, rejection records, leaderboard reports, and iterative quality-control rules. This gives me strong experience in high-precision AI evaluation tasks where the output must be judged not only by language quality, but also by mathematical validity, domain logic, and downstream performance.

2026 - Present

LLM Response Evaluation & RLHF Data Review

TextTextRLHFRLHF

Evaluated AI-generated responses for accuracy, relevance, completeness, reasoning quality, safety, and instruction-following. Compared multiple model outputs, identified hallucinations, weak reasoning, factual errors, and low-quality answers, and provided structured feedback to improve AI training data quality. Strong focus on Chinese-language tasks, bilingual Chinese-English review, and high-consistency judgment standards.

2025 - Present

AI Evaluation & Automation Tooling (Independent projects)

Other

Built AI evaluation and automation workflows that score content quality and output readiness for multiple criteria. Designed rubric-like scoring processes for SEO/GEO readiness, technical recommendations, and channel-specific assessment. Established systematic tracking with structured records, metric monitoring, pass/fail gates, and failure logging to drive iterative improvements. • Implemented scoring and quant-research automation habits for candidate generation and metric tracking. • Used structured spreadsheets and JSON/JSONL-style records for error labeling and reproducible evaluation. • Added pass/fail gates and leaderboard outputs to monitor quality over iterations. • Maintained failure logs and iterative refinement loops to improve labeling/evaluation consistency.

2025 - Present

Multimodal Personal Knowledge / AI Workflow Design (Independent product exploration)

DocumentDocumentQuestion AnsweringQuestion Answering

Designed multimodal AI knowledge workflows that support labeling-like tasks across PDF/Word/image/video/audio ingestion and OCR/ASR extraction. Implemented pipeline concepts for classification, tagging, long-term memory, and retrieval-augmented question answering with traceable sources. Added quality control and human-review checkpoints to validate outputs against source material and flag unsupported claims. • Covered ingestion and extraction across multiple modalities (documents, images, audio/video). • Specified OCR/ASR extraction and subsequent classification/tagging and retrieval steps. • Implemented traceable source retrieval and quality control for AI-generated answers. • Defined inspection and failure handling via human-review gates for factuality and supportability.

2025 - Present

Biomedical Image Segmentation Research (Academic / independent)

Don't disclose

Conducted research-focused labeling QA for biomedical image segmentation, emphasizing when specific evaluation metrics are appropriate. Applied segmentation review logic to lung CT scenarios including lung region segmentation and tumor/nodule plus organ-at-risk concepts. Used metric-driven judgment to connect model outputs with clinical relevance and common failure modes. • Compared Dice, IoU, and Hausdorff Distance with sensitivity, recall, and false-negative rate. • Reviewed surface/boundary accuracy and boundary-error behavior for segmentation quality. • Analyzed medical-AI limitations impacting annotation and evaluation alignment (e.g., class imbalance, overfitting, privacy). • Mapped metric-clinical mismatch risks to practical dataset/labeling considerations for CT tasks.

2024 - Present

Education

U

University of Edinburgh

Master of Science, Communications and Signal Processing

Master of Science
2024 - 2025
N

Nanjing University of Aeronautics and Astronautics

Bachelor of Science, Electronic Information Science and Technology

Bachelor of Science
2020 - 2024

Work History

Z

ZHISHU

Chief AI Trainer

Nanjing
2025 - Present