Senior Python Engineer (AI testing & evaluation automation) | LinkedIn | 2020–2024
Built functional black-box tests to evaluate Python and Rust CLI tools, supporting systematic AI testing workflows. Automated repetitive test-case generation using LLM assistance to improve iteration speed and reliability of evaluation. Managed reproducible, secure Docker environments to ensure consistent test execution across Linux platforms and teams. • Created automated coverage targets (e.g., 95%) via functional test suites • Standardized testing practices and improved pipeline efficiency • Reduced development time by ~30% using LLM-generated test cases • Collaborated cross-functionally to align AI evaluation approaches