LLM Evaluation & Repository Validation Engineer (Independent Contractor) — Habitat Inc.
Contributed to designing and validating AI coding assistant benchmark tasks by reviewing AI-generated code against reference solutions. Created and refined multi-language test cases to probe edge cases, specification adherence, and build/runtime behavior for model evaluation. Delivered transparent, constructive feedback to distinguish ambiguous prompts from genuine model reasoning limitations. • Compared AI-generated implementations to golden/reference answers and identified subtle logical errors • Authored diverse, challenging test scenarios derived from real-world open-source commits (Python, Go, Java) • Documented findings, best practices, and recommendations to improve benchmark dataset quality • Helped define coding standards and task acceptance benchmarks targeting >0% to ≤50% AI success rates