Build secure sandboxes, evaluation pipelines, and developer tooling for LLM agent training as a remote Python infrastructure engineer. Work 20+ hours weekly from the US or Canada at experience-based rates of $34 to $42 per hour.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build their profiles, and apply in minutes.
- Remote contractor opportunity for talent located in the United States or Canada
- Part-time schedule requiring 20+ hours per week
- Hourly compensation listed by experience tier: Junior $34, Middle $37, or Senior $42 USD
About AI Training Infrastructure
AI training is the human side of building modern artificial intelligence. Engineers and specialists create the sandboxes, task frameworks, testing systems, and scoring pipelines that help researchers evaluate how AI agents perform in realistic environments.
This work supports a fast-growing field where contributors help shape cutting-edge AI systems. It is remote and can offer flexible work arrangements for people who want to build experience in the industry.
- Support LLM training and agent evaluation workflows
- Create reliable tools that let AI researchers iterate quickly
- Contribute to systems used for coding-focused AI tasks
The Role
OpenTrain AI is seeking a senior-minded Python Infrastructure Engineer to own infrastructure underpinning LLM training workflows. You will build secure sandboxes, reusable task frameworks, scoring pipelines, automated evaluation systems, and modern developer environments for agent evaluation.
The role combines production Python engineering with Linux, Docker, CI/CD, backend services, testing, and security. Clear English communication and a collaborative approach are essential because you will guide AI researchers through the tools you build.
- Work remotely from the United States or Canada
- Contribute 20+ hours per week
- Work as a contractor on a part-time engagement
- Use Python to build infrastructure for LLM and agent evaluation
What You'll Do
You will deliver reusable repositories and dependable engineering systems that make agent-task development and evaluation faster, safer, and easier to repeat. The work includes both hands-on implementation and collaboration with researchers.
- Build secure sandboxes and task frameworks for AI agents
- Develop scoring pipelines and automated evaluation workflows
- Create modular REST or asynchronous services with FastAPI or Flask
- Design development environments using devcontainers, Makefiles, .env workflows, and pre-commit hooks
- Create GitHub Actions or comparable CI/CD pipelines that lint, test, build, and deploy
- Write unit, integration, and functional tests with pytest
- Harden Docker images and integrate security scanners such as Trivy or Snyk in CI
- Support researchers through concise documentation, pair programming, and clear communication
- Use AI coding assistants such as Cursor, Claude Code, or Copilot responsibly
Required Qualifications
You should bring at least five years of professional Python experience and be comfortable owning production-grade infrastructure. A computer science or engineering degree is preferred, or you may qualify through equivalent hands-on experience.
- 5+ years of professional Python experience, including async I/O, packaging, and refactoring
- Strong Linux skills, including bash, grep, curl, jq, permissions, and basic networking
- Deep Docker experience with multi-stage Dockerfiles, image optimization, and docker-compose
- Experience designing CI/CD pipelines with GitHub Actions or similar tools
- Proficiency with FastAPI or Flask, Pydantic validation, and structured logging
- Test-driven development experience with pytest and high-coverage targets
- Experience creating sandboxes, scoring pipelines, or evaluation frameworks for AI agents
- Security awareness, including least-privilege practices and container hardening
- Strong version-control discipline, including semantic commits, branch hygiene, and thorough code reviews
- Clear English communication and the ability to collaborate with and guide AI researchers
Preferred Experience and Work Style
This opportunity suits engineers who can work independently while collaborating closely with technical researchers. You should be comfortable learning within a rapidly evolving AI environment and turning complex infrastructure needs into maintainable tools.
- Prior experience training or evaluating AI systems, especially coding-focused tasks
- Experience with Kubernetes is a plus
- Hands-on use of AI coding assistants
- A collaborative mindset and effective pair-programming habits
- Concise technical documentation and practical researcher support
- Senior-minded ownership, even within an intermediate experience profile
Application and Screening Process
Short-listed candidates will complete a timed HackerRank assessment and a platform coding test before progressing to recruiter interviews. Candidates must be able to complete both assessments within 48 hours of receiving an invitation.
- Submit your application through OpenTrain
- Complete the timed HackerRank assessment if shortlisted
- Complete the platform coding test if shortlisted
- Progress to recruiter interviews after the assessments
- Be ready to complete the screening steps within 48 hours of invitation