Skip to content
OpenTrain AIFor AI Companies

Python Infrastructure Engineer, LLM Tooling & Agent Evaluation

Build and own the infrastructure that powers LLM training and agent evaluation at OpenTrain AI; remote for US & Canada, 20+ hours/week, contractor roles with hourly pay tiers (Junior $34 / Middle $37 / Senior $42). Shortlisted candidates complete a timed HackerRank and platform coding test.

OpenTrain AI

Coding & Software

100% Remote Hourly · $37/hr

$37/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Jul 25, 2025

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain AI

OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We help people start and grow careers teaching AI by connecting contributors with meaningful projects, and we hire engineers and contractors to build the tools that make high-quality human-in-the-loop training possible.

This role is a hands-on engineering position inside OpenTrain AI: you'll develop the developer-facing infrastructure and pipelines that let researchers and labelers iterate quickly and safely on LLMs and agents.

About AI training and this work

AI training (data labeling, annotation, and human feedback work) is the human side of building modern AI. Engineers who build training infrastructure enable people to create, evaluate, and refine examples that state-of-the-art models learn from.

Contributors in this space enjoy flexible, remote work and make an outsized impact on how AI systems behave. This role focuses on the tooling that lets researchers run repeatable, secure, and automatable agent evaluations.

The role

We're seeking senior-minded Python engineers to own the infrastructure underpinning our LLM-training workflow and agent evaluation tooling. You will design and implement secure sandboxes, task frameworks, scoring pipelines, and developer environments that let experts iterate rapidly and safely.

This is a contractor, part-time role at OpenTrain AI. Positions are fully remote for talent located in the United States and Canada and typically require 20+ hours per week.

What you'll deliver

You will ship reusable repositories, automated evaluation and scoring pipelines, and polished developer environments that researchers and labelers can adopt immediately. Your work should be secure, test-driven, and easy to operate.

  • Secure sandboxes and isolated runtimes for executing agent tasks and tests.
  • Task frameworks and harnesses for defining, running, and reproducing agent evaluations.
  • Automated scoring pipelines that produce reliable metrics and artifacts for analysis.
  • Developer environments: devcontainers, Makefiles, .env workflows, pre-commit hooks, and clear setup docs.
  • CI/CD pipelines (GitHub Actions preferred) that lint, test, build, cache, and deploy artifacts safely.

Key requirements

You must be able to work autonomously, communicate clearly with researchers, and produce production-grade code and tooling that other engineers and non-engineer researchers can use and maintain.

  • 5+ years professional Python experience: production-grade code, async I/O, packaging, and refactoring.
  • CS/Engineering degree or equivalent hands-on experience.
  • Test-driven mindset: writes unit, integration, and functional tests with pytest; targets high coverage.
  • Linux power-user: bash, grep, curl, jq, permissions, basic networking.
  • Docker expertise: multi-stage Dockerfiles, image optimization, docker-compose (Kubernetes a plus).
  • CI/CD ownership: designs GitHub Actions or similar; manages secrets, caching, and secure workflows.
  • FastAPI or Flask proficiency: modular REST/async services with Pydantic validation and structured logging.
  • Dev-environment setup: creates devcontainers, Makefiles, .env workflows, and pre-commit hooks.
  • LLM/agent infrastructure exposure: built sandboxes, scoring pipelines, or evaluation frameworks for agents.
  • Experience using AI coding assistants responsibly (Copilot, Claude Code, Cursor, etc.).
  • Strong collaboration and communication: supports researchers, writes concise docs, and pair-programs effectively.
  • Security awareness: least-privilege, Docker image hardening, integrates scanners (Trivy/Snyk) in CI.
  • Version-control discipline: semantic commits, branch hygiene, and thorough code reviews.
  • Screening readiness: able to complete a timed HackerRank assessment and a platform coding test within 48 hours of invite.

Who should apply

Apply if you enjoy building developer-facing infrastructure, have deep Python and Linux skills, and want to enable safe, repeatable LLM and agent evaluations. This role suits engineers who pair well with researchers and value testability, security, and reproducibility.

Prior experience training or evaluating AI systems on OpenTrain, especially coding-focused tasks, is a strong bonus and will help you ramp faster.

Compensation, schedule, and hiring process

This is a contractor, part-time engagement requiring 20+ hours per week and is open to candidates based in the United States and Canada.

Hourly pay is tiered by experience: Junior $34/hr, Middle $37/hr, and Senior $42/hr (USD). Shortlisted candidates will complete a timed HackerRank assessment and a platform coding test before recruiter interviews.

  • Employment type: Contractor, Part-time.
  • Time commitment: 20+ hours/week.
  • Location: Remote (United States and Canada only).
  • Assessment: HackerRank + platform coding test; must be able to complete within 48 hours of invite.