For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
Y

Yifei H.

Multimodal LLM Cluster Deployment & Document Extraction Pipeline — AI Infrastructure Specialist

China flagN/A, China

Key Skills

Software

No software listed

Top Subject Matter

Multimodal document extraction and data curation for LLM/RAG pipelines
AI agent orchestration and preparation of dataset-driven LLM outputs

Top Data Types

DocumentDocument
TextText
ImageImage

Top Task Types

Data CollectionData Collection
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)

Freelancer Overview

Multimodal LLM Cluster Deployment & Document Extraction Pipeline — AI Infrastructure Specialist. Brings 8+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include vLLM and sglang and Dify and LangChain. AI-training focus includes data types such as Document and Text and labeling workflows including Data Collection and Prompt + Response Writing (SFT).

Labeling Experience

LLM-Powered Data Analysis Agent Platform — AI Engineer & Full-Stack Developer

TextTextPrompt + Response Writing (SFT)Prompt + Response Writing (SFT)

Developed an LLM-powered conversational data analysis agent that ingests user-submitted datasets, parses complex schemas, and performs multi-step analytical reasoning loops. Implemented Retrieval-Augmented Generation (RAG) with vector databases to apply domain-specific contextual constraints. Deployed and managed Dify to convert fragile AI scripts into scalable, reproducible enterprise workflows for iterative AI output generation. • Built agentic data analysis workflows with schema parsing • Applied RAG using vector databases for contextual constraint enforcement • Operationalized workflows in Dify for reproducible enterprise runs • Integrated multi-step reasoning loops for automated analysis responses

Present

Multimodal LLM Cluster Deployment & Document Extraction Pipeline — AI Infrastructure Specialist

DocumentDocumentData CollectionData Collection

Built and operated an automated document parsing pipeline using multimodal OCR and visual reasoning to convert unstructured, heavily stylized documents into structured fields. Collected and curated large corpora of text and media via distributed scraping, followed by sanitization to produce high-quality datasets for downstream knowledge bases. This work supported reliable data preparation for generative AI/RAG applications. • Automated extraction of structured data fields from unstructured documents • Configured multimodal model workflows for vision-OCR and reasoning • Implemented distributed scraping for corpus aggregation and cleansing • Prepared cleaned datasets for downstream RAG knowledge bases and fine-tuning

Present

Work History

C

Connected Vehicle Cloud Service

Core Data Engineer

N/A
2022 - 2024
L

LLM-Powered Data Analysis Agent Platform

AI Engineer and Full-Stack Developer

N/A
2020 - 2022