Full-Stack Senior Engineer and AI Evaluation, Upwork & Toptal (Freelance) (01/2021–Present)
Built fully automated assessment systems to evaluate the accuracy and quality of AI code models at scale. Implemented CI/CD with in-between model testing to support reproducible research workflows and continuous evaluation. Developed infrastructure and distributed evaluation pipelines to stress-test AI-assisted codification and measure performance metrics. • Scaled benchmarking pipelines to handle millions of generated code outputs • Created full-stack dashboards to visualize AI performance metrics and error patterns • Implemented distributed evaluation using Docker and Kubernetes clusters • Built safe APIs and back-end services for large-scale experimentation pipelines