AI Coding Agent Evaluator (Freelance) — Schema Integrity & Vectorized Inference testing for ML serving pipelines
Evaluated reliability of LLM-generated ML serving pipelines via adversarial schema validation and inference-path testing to identify boundary failures and runtime crashes. Authored structured PASS/FAIL reports and analyzed implicit constraint reasoning gaps and missing domain-level test coverage. Designed controlled test harnesses that reproduce validation failures across row-oriented and column-oriented API contracts. • Performed schema integrity testing for Pydantic V2 validation paths • Tracked regression cases where categorical inputs bypass validation and crash ONNX inference at runtime • Created boundary-focused test harnesses for different API contract orientations • Produced structured evaluation outputs (PASS/FAIL) for validation outcomes and failures