Skip to content
OpenTrain AIFor AI Companies

DevOps Engineer, GPU And LLM Infrastructure

Apply now

DevOps Engineer, GPU And LLM Infrastructure

Work remotely as a contractor building GPU infrastructure and production LLM serving systems on Google Cloud. Join OpenTrain’s growing AI-training ecosystem and apply your DevOps expertise to cutting-edge model deployment.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Sep 11, 2026

Posted

Open worldwide

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring for this remote contractor engagement, helping technical professionals discover opportunities, build a credible AI work profile, and apply in minutes.

  • Remote contractor engagement
  • Part-time schedule of 20+ hours per week
  • Worldwide opportunity
  • English-language role

About AI Training Work

AI training is the human side of building artificial intelligence. Behind modern models are technical systems that provision computing resources, manage data pipelines, serve model outputs, and keep production workloads reliable. This fast-growing field gives contributors the opportunity to work with cutting-edge AI systems from anywhere.

  • Support the infrastructure behind AI model development and deployment
  • Work with production large language model serving systems
  • Contribute to a rapidly evolving technical industry

The Role

OpenTrain is recruiting a DevOps Engineer to build and operate GPU infrastructure and production LLM serving systems on Google Cloud Platform. You will support AI model development and deployment through GPU provisioning, container orchestration, scalable inference, data pipelines, observability, reliability, and cost optimization.

This is an entry-level contractor opportunity requiring 20+ hours per week. The role calls for strong production infrastructure experience and hands-on expertise with Google Cloud, GPU operations, and LLM serving technologies.

  • Build and operate infrastructure on Google Cloud Platform
  • Support scalable inference and production model serving
  • Help maintain reliable, observable, and cost-conscious AI systems

What You'll Do

You will provision and manage GPU infrastructure across GKE, Compute Engine, Cloud Run, and related Google Cloud services. You will also deploy, tune, and operate production LLM serving stacks using vLLM, Triton Inference Server, and Hugging Face Transformers.

The work includes containerized and serverless deployments, infrastructure automation, machine learning data pipelines, and operational improvements for GPU fleets and production inference services.

  • Provision and manage GPU infrastructure on GKE, Compute Engine, Cloud Run, and related Google Cloud services
  • Deploy, tune, and operate LLM serving stacks with vLLM, Triton Inference Server, and Hugging Face Transformers
  • Build containerized and serverless deployments with Docker, CI/CD, and Infrastructure as Code
  • Design and maintain reliable data pipelines for machine learning and inference workloads
  • Implement observability, autoscaling, reliability, and cost optimization for GPU fleets and inference services

Required Skills

You should have strong Python programming skills and hands-on experience building or operating production infrastructure. Practical Google Cloud Platform experience is required, along with the ability to manage GPU environments and serve large language models in production.

  • Strong Python programming and production infrastructure engineering experience
  • Practical GCP experience with GKE, Cloud Run, and related services such as GCS, Pub/Sub, or Cloud Functions
  • Hands-on GPU infrastructure operations, including CUDA, drivers, GPU node pools, quotas, and autoscaling
  • Production LLM serving experience with vLLM or Triton Inference Server
  • Strong Docker and container orchestration fundamentals
  • Experience building reliable ML data pipelines with Dataflow, Airflow, Cloud Composer, or similar tools

Helpful Background

The following experience will help you succeed in the role and extend your ability to operate AI infrastructure at production scale.

  • Hugging Face Transformers
  • Serverless deployments
  • Infrastructure as Code
  • CI/CD
  • Observability
  • Production cost optimization

Why Work With OpenTrain

OpenTrain helps professionals build careers in AI training and data labeling, an industry that includes model evaluation, data preparation, software and infrastructure support, and other work that shapes how AI systems are built. Creating an OpenTrain account is free, and a stronger profile can help you showcase relevant experience and grow a long-term AI portfolio.

  • Find and apply for AI-focused work in one place
  • Build a profile around your technical experience
  • Work remotely with flexible project opportunities
  • Contribute to state-of-the-art AI systems

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

GPU Programming Software Engineer

Build and optimize GPU software for LLM training tasks using CUDA, WebGPU, GLSL, and C++. This expert contractor role offers worldwide, part-time work at $60-$95 per hour.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Expert level
Hourly · $60–$95/hr

Posted Jul 29, 2026

Python Infrastructure Engineer for LLM Training

Build secure sandboxes, evaluation pipelines, and developer tooling for LLM agent training as a remote Python infrastructure engineer. Work 20+ hours weekly from the US or Canada at experience-based rates of $34 to $42 per hour.

Coding & Software
Computer Code Programming
Remote · Canada, United States
English
Part-time · Flexible
Intermediate level
Hourly · $37/hr

Posted Jul 25, 2025

Python Infrastructure Engineer, LLM Agent Tooling

Build the secure sandboxes, developer environments, CI/CD pipelines, and scoring systems behind LLM training and agent evaluation. This remote contractor role offers 20+ hours per week for Python engineers in eligible Asia-Low countries.

Coding & Software
Computer Code Programming
Remote · Bangladesh, Georgia, India +7 more
English
Part-time · Flexible
Intermediate level
Hourly · $12.5/hr

Posted Jul 25, 2025