Own production infrastructure and backend reliability for AI evaluation workflows in a remote three-month contractor role. Apply through OpenTrain to support Cloud Run, Python services, Kubernetes workloads, and high-volume production operations.
About OpenTrain
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional profile, and grow in this fast-moving field. Creating an OpenTrain account is free.
About AI Training Infrastructure
AI training and evaluation depend on reliable technical systems as well as human expertise. Engineers help keep the services, containers, databases, and deployment workflows running so AI evaluation work can be completed consistently at scale.
This role supports the infrastructure behind cutting-edge AI workflows. It is a chance to contribute to a rapidly growing industry while working remotely on production systems that handle demanding, concurrent workloads.
The Role
OpenTrain is recruiting a Senior GCP Infrastructure and Backend Engineer to operate production services supporting AI evaluation and training workflows. You will own deployments, reliability, uptime, incident response, and technical support for systems that run and package high-volume workloads.
The work includes maintaining quality-control services, supporting Cloud Run infrastructure, and resolving issues across concurrent runs and large container fleets. This is a remote contractor engagement expected to last three months.
- Remote contractor engagement lasting approximately three months
- Eight hours per day, including six to eight hours of overlap with Pacific Time
- The structured role details specify a commitment of 20+ hours per week
- Candidates based in the United States or Latin American time zones are preferred
- Worldwide candidates may apply; English is the listed working language
What You’ll Do
You will work across production operations, backend services, and infrastructure reliability. The role requires independent troubleshooting and close attention to service performance, deployment health, and operational continuity.
- Manage and monitor production deployments across Google Cloud services
- Maintain reliability, uptime, and performance across distributed workloads
- Debug backend, infrastructure, and deployment issues in production
- Support Cloud Run runner and final-packaging systems
- Maintain and improve quality-control services used in AI evaluation workflows
- Troubleshoot failures across hundreds of concurrent runs and large pod fleets
- Respond independently to trainers’ technical questions and operational requests
- Collaborate on incident resolution, service improvements, and operational documentation
Required Experience And Technical Skills
This role requires five or more years of hands-on experience in infrastructure engineering, backend engineering, or production systems, along with the ability to manage production responsibilities independently.
- Five or more years of hands-on infrastructure, backend, or production-systems experience
- Strong production experience with Google Cloud Platform, especially Cloud Run and related services
- Advanced Python backend development skills, including debugging and maintaining production services
- Strong knowledge of Kubernetes, containerized workloads, SQL databases, and distributed systems
- Experience operating highly concurrent production workloads and large pod fleets
- Ability to independently manage deployments, incidents, uptime, bug fixes, and trainer technical requests
Who Should Apply
Apply if you are comfortable taking ownership of production systems and solving infrastructure, deployment, and backend problems without constant supervision. The strongest fit will combine advanced Python skills with practical GCP operations experience and a clear understanding of distributed, containerized workloads.
- Senior infrastructure or backend engineers with production ownership experience
- Engineers experienced with Cloud Run, Kubernetes, and large-scale container operations
- Candidates who can overlap substantially with Pacific Time
- Professionals interested in applying their technical expertise to AI evaluation and training systems
How To Apply Through OpenTrain
Create or update your free OpenTrain profile to present your experience and apply for this opportunity. OpenTrain helps contributors build a credible portfolio in AI training and data labeling while finding projects that match their skills.
Review the role requirements carefully, highlight your GCP and Python production experience, and submit your application through OpenTrain.