For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
C
Chen J.

Chen J.

Lead Developer & AI Evaluator | Modular RAG MCP Server

China flagZhuhai, China

Key Skills

Software

No software listed

Top Subject Matter

LLM evaluation for RAG retrieval quality and hallucination mitigation
LLM-generated code evaluation (correctness, security/safety, clarity)
Legal Services & Contract Review

Top Data Types

TextText
DocumentDocument

Top Task Types

RLHFRLHF
Entity (NER) ClassificationEntity (NER) Classification

Freelancer Overview

Lead Developer & AI Evaluator | Modular RAG MCP Server. Brings 2+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Internal and Proprietary Tooling. Education includes Bachelor of Engineering, Zhuhai University of Science and Technology (2025). AI-training focus includes data types such as Text, Computer Code, and Programming and labeling workflows including Evaluation, Rating, and Computer Programming.

Labeling Experience

Data Annotation & AI Evaluation Specialist - OpenTrainAI

TextTextRLHFRLHFEntity (NER) ClassificationEntity (NER) Classification

You served as a Data Annotation & AI Evaluation Specialist candidate focused on improving LLM response quality through evaluation-oriented engineering. You applied structured evaluation thinking to assess retrieval-augmented outputs, including relevance, faithfulness, and context recall. Your work emphasized systematic rubric design and quality measurement backed by practical Python and data pipeline skills. • Evaluated LLM outputs with quality and safety dimensions • Designed and applied structured annotation/evaluation rubrics • Built and validated retrieval-quality metrics for RAG systems • Used Python, SQL, and data processing pipelines to support evaluation workflows.

2024 - 2025

Lead Developer & AI Evaluator | Modular RAG MCP Server

TextText

Led development and evaluation of a Modular RAG MCP Server to assess LLM outputs using systematic retrieval-quality and response-quality metrics. Designed a two-stage retrieval pipeline and incorporated Ragas evaluation for retrieval faithfulness, answer relevance, and context recall. Produced structured evaluation workflows to support annotation-style quality assessment for retrieval-augmented generation.• Evaluated LLM outputs against retrieval-faithfulness and context-recall criteria• Tuned retrieval (BM25 + dense hybrid with RRF fusion) and fine ranking (cross-encoder)• Implemented a provider-switching backend to standardize evaluation across LLM/embedding/reranker variants• Built Streamlit observability dashboards to monitor evaluation signals end-to-end

2024 - 2025

Developer & Evaluator (AI Coding Assistant) - Ai-Coding-Assistant

TextTextRLHFRLHFEntity (NER) ClassificationEntity (NER) Classification

You developed an AI coding assistant supporting multi-model code generation and evaluation with structured safety checks. You built line-by-line explanation and error analysis flows while enabling one-click switching across multiple LLM providers. You also designed a layered code safety evaluation framework using AST analysis, whitelisting, subprocess isolation, and timeout termination. • Implemented multi-model code generation and explanation workflows • Created AST-based dangerous-call interception and safe-call whitelisting • Added subprocess isolation and timeout termination for risk control • Produced annotated evaluation feedback for correctness, security, and clarity.

2024 - 2024

Developer & Evaluator | AI Coding Assistant (bliss-fox/ai-coding-assistant)

Built an AI coding assistant that generated and then evaluated code with structured, rubric-based assessment focused on correctness, safety, and explanation quality. Implemented an AST-based four-layer safety evaluation framework and produced annotated feedback reports from evaluation results. Applied systematic evaluation to compare multiple LLM providers for coding task performance and clarity.• Intercepted dangerous AST calls and enforced a built-in function whitelist• Used subprocess isolation and timeout termination to reduce execution risk• Assessed generated code for correctness and security risks and generated annotation-style feedback• Evaluated code explanation quality and task compliance across OpenAI, DeepSeek, and Qwen providers

2024 - 2024

Education

Z

Zhuhai University of Science and Technology

Bachelor of Engineering, Computer Science

Bachelor of Engineering
2021 - 2025

Work History

M

Modular RAG MCP Server

Lead Developer & AI Evaluator (RAG Systems)

Zhuhai
2024 - 2025
O

OpenTrainAI

Data Annotation & AI Evaluation Specialist

Zhuhai
2024 - 2025