Lead AI Evaluator & Content Strategist
- Model Testing & Hallucination Detection: Conducted rigorous stress tests on leading LLMs (ChatGPT, Claude, Gemini) to identify performance boundaries and logical vulnerabilities. Produced 26 deep-dive evaluation reports focusing on model reasoning and integration. - Intent Annotation & Instruction Optimization: Diagnosed and optimized thousands of real-world user prompts within a high-net-worth AI community. Successfully performed intent recognition tasks to bridge the gap between human requirements and model execution. - Structured Dataset Synthesis: Extracted core logic from unstructured AI industry data to create high-quality training corpora. Expertise in applying industry-standard labeling methods to improve model helpfulness and accuracy.