AI Evaluation & Software Engineer (Contractor)
As an AI Evaluation & Software Engineer at Habitat Code, I developed and curated complex coding challenges to evaluate the reasoning and coding capabilities of frontier AI models. My responsibilities included creating high-fidelity datasets by engineering evaluation benchmarks using real-world software defects. I ensured the accuracy and challenge level of these benchmarks to facilitate the effective training, evaluation, and improvement of AI agents. • Identified, selected, and adapted high-complexity coding scenarios from public repositories. • Authored 'golden patch' solutions and designed automated test suites for rigorous model evaluation. • Focused on low-pass-rate problem sets to specifically target AI model weaknesses and blind spots. • Maintained reproducible environments and high-quality data pipelines for AI training and quality assurance.