Skip to content
OpenTrain AIFor AI Companies

DevOps Engineer, GPU and LLM Infrastructure

Build and operate GPU infrastructure and production LLM serving systems on Google Cloud Platform. This remote, two-month contractor role includes an 8-hour daily commitment and Pacific Time overlap.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Sep 4, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI recruits contractors for specialized projects where technical professionals help build the infrastructure, systems, and evaluation workflows behind modern artificial intelligence.

Create a free OpenTrain account to build a profile, showcase relevant experience, discover matching opportunities, and apply in minutes.

About AI Infrastructure Work

AI training depends on reliable infrastructure for model development, evaluation, inference, and data processing. Engineers working in this field help operate the cloud platforms, GPU fleets, deployment systems, and pipelines that support increasingly capable AI models.

This role focuses on the technical foundation behind scalable model inference, including GPU provisioning, container orchestration, observability, reliability, autoscaling, and cost optimization.

The Role

OpenTrain AI is recruiting a DevOps Engineer to build and operate GPU infrastructure and production LLM serving systems on Google Cloud Platform. You will support AI systems and data pipelines used by teams working on foundation models, LLM evaluation, and reasoning.

This is a remote contractor engagement lasting two months. The listed commitment is 8 hours per day, including a 4-hour overlap with Pacific Time. Compensation details are handled through standard OpenTrain project defaults.

  • Remote contractor opportunity
  • Two-month engagement
  • 8 hours per day with 4 hours of Pacific Time overlap
  • 20+ hours per week listed in the project requirements
  • English-language role
  • Worldwide eligibility

What You'll Do

You will provision, deploy, tune, and operate cloud infrastructure and production inference services for machine learning workloads. The work combines hands-on platform engineering with reliability, data pipeline, and cost-efficiency improvements.

  • Provision and manage GPU infrastructure on GCP using GKE, Compute Engine, Cloud Run, and related Google Cloud services.
  • Deploy, tune, and operate production LLM serving stacks with vLLM, Triton Inference Server, and Hugging Face Transformers.
  • Build containerized and serverless deployments using Docker, CI/CD, and infrastructure as code.
  • Design and maintain reliable data pipelines supporting machine learning and inference workloads.
  • Improve observability, autoscaling, reliability, and cost efficiency for GPU fleets and production inference services.

Required Skills

This opportunity is listed at the entry level, but the project requires practical experience operating production infrastructure and specialized GPU and LLM systems. Candidates should be able to work hands-on with the tools and services below.

  • Strong Python programming for production infrastructure.
  • Practical Google Cloud Platform experience, including GKE, Cloud Run, GCS, Pub/Sub, Cloud Functions, or equivalent services.
  • Hands-on experience provisioning and operating GPU infrastructure, including CUDA, drivers, GPU node pools, quotas, and autoscaling.
  • Production LLM serving experience with vLLM or Triton Inference Server.
  • Strong Docker and container orchestration fundamentals.
  • Experience building machine learning data pipelines with Dataflow, Airflow, Cloud Composer, or similar technologies.

Helpful Background

The following experience is useful for supporting the infrastructure and deployment needs of this engagement.

  • Hugging Face Transformers
  • Infrastructure as code
  • CI/CD systems
  • Serverless deployments
  • Production observability

Why Build Your AI Career With OpenTrain

AI training and data labeling are the human side of building artificial intelligence. Technical contributors help shape how modern AI systems are developed, evaluated, and operated, making this a fast-growing area for people with cloud, software, and machine learning infrastructure skills.

OpenTrain gives you one place to build a credible AI-work profile, manage opportunities, and grow toward a durable career in this cutting-edge industry. Apply through OpenTrain to get started.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

GPU Programming Software Engineer

Build and optimize GPU software for LLM training tasks using CUDA, WebGPU, GLSL, and C++. This expert-level contractor role offers worldwide remote work at $60 to $85 per hour for 20+ hours weekly.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Expert level
Hourly · $60–$85/hr

Posted Jul 29, 2026

Python Infrastructure Engineer for LLM Training

Build secure sandboxes, evaluation pipelines, and developer tooling for LLM agent training as a remote Python infrastructure engineer. Work 20+ hours weekly from the US or Canada at experience-based rates of $34 to $42 per hour.

Coding & Software
Computer Code Programming
Remote · Canada, United States
English
Part-time · Flexible
Intermediate level
Hourly · $37/hr

Posted Jul 25, 2025

Python Infrastructure Engineer, LLM Agent Tooling

Build the secure sandboxes, developer environments, CI/CD pipelines, and scoring systems behind LLM training and agent evaluation. This remote contractor role offers 20+ hours per week for Python engineers in eligible Asia-Low countries.

Coding & Software
Computer Code Programming
Remote · Bangladesh, Georgia, India +7 more
English
Part-time · Flexible
Intermediate level
Hourly · $12.5/hr

Posted Jul 25, 2025