Multi-Model Blind A/B Testing & Logical Error Auditing
Conducted an independent, structured evaluation project focused on stress-testing the logical boundaries and constraint adherence of frontier LLMs (including Gemini and Claude). Designed and executed a series of 50+ bilingual (English/Chinese) prompts using structured prompting techniques like Chain-of-Thought (CoT) and negative constraints to evaluate multi-step reasoning capabilities. My hands-on work involved running blind A/B preference testing, meticulously cataloging model outputs based on logical consistency, information density, and strict adherence to formatting instructions. I maintained a granular tracking log to analyze and categorize text hallucinations, specifically capturing source misattributions and logical overgeneralizations in technical summaries. For every evaluation, I drafted objective, evidence-based justifications in professional English, breaking down the exact mathematical or logical failures of the underperforming model. This project demonstrated my ability to apply zero-variance operational discipline—honed during my career managing complex maritime data logs—to high-density AI data auditing.