Senior Python Infrastructure Engineer for LLM Tooling
Join OpenTrain AI to build the Python infrastructure that powers LLM training and agent evaluation. Fully remote for Asia‑Low region, 20+ hrs/week contract work with tiered hourly pay; requires 5+ years Python, Docker, CI/CD, FastAPI/Flask, and Linux expertise.
Coding & Software
100% Remote Hourly · $12.5/hr
$12.5/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 25, 2025
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for building careers in AI training and data labeling. We help people discover projects, grow skills, and contribute directly to how modern AI systems are trained and evaluated.
Working with OpenTrain means joining a fast‑growing industry where human contributors shape AI behavior, enjoy flexible remote schedules, and gain experience on real LLM and agent evaluation pipelines.
Why AI‑training infrastructure matters
AI training is the human side of building intelligent systems: engineers and annotators prepare, test, and evaluate data and models. Reliable developer tooling, secure sandboxes, and reproducible scoring pipelines are critical to reproducible, safe, and efficient LLM research and evaluation.
Your work makes researcher workflows repeatable, secure, and faster.
Good infra reduces risk and improves the quality of model evaluations.
The role
We need senior‑minded Python engineers to design and maintain infrastructure that powers LLM training, agent tooling, and evaluation workflows. This is a contract, part‑time role (20+ hours/week) open to candidates located in the Asia‑Low region (Afghanistan through Vietnam).
Employment type: Contractor, Part‑time.
Time requirement: 20+ hours per week.
Location: Fully remote for Asia‑Low region.
What you'll build and support
You will create reusable repositories, developer environments, secure sandboxes, scoring pipelines, and CI/CD that researchers use to evaluate agent performance. Expect a mix of hands‑on coding, design, and collaboration with research teams.
Design and maintain secure sandboxes and task environments for LLM/agent evaluation.
Build and maintain CI/CD pipelines (GitHub Actions or similar) that lint, test, build, and deploy.
Containerize services with multi‑stage Dockerfiles and docker‑compose; K8s experience is a plus.
Author test suites with pytest and enforce test coverage and reliability.
Integrate security scanners (Trivy/Snyk) and apply least‑privilege principles in images and CI.
Requirements & qualifications
You must bring production experience and a testing mindset: write clean, test‑driven Python, operate confidently on Linux, and own CI/CD. Selected candidates will complete a timed HackerRank assessment and a platform coding test before interviews.
5+ years professional Python — production code, packaging, async I/O, refactoring legacy modules.
Testing: writes unit/integration/functional tests with pytest; focuses on coverage and reliability.
Linux power‑user: bash, grep, curl, jq, systemd; basic networking and permissions.
Container expertise: multi‑stage Dockerfiles, docker‑compose; Kubernetes is a plus.
FastAPI or Flask proficiency: modular REST/async services, auth, validation with Pydantic, logging.
Dev env and tooling: devcontainer.json, Makefiles, .env workflows, pre‑commit, PR templates.
LLM/agent infra exposure: experience building sandboxes, scoring pipelines, or evaluation frameworks.
Familiarity with AI coding assistants (Cursor, Claude Code, Copilot, etc.) and safe usage practices.
Collaboration: mentors teammates, pairs with researchers, and writes concise docs.
Security awareness: hardens images, integrates scanners in CI, practices least‑privilege.
Compensation, assessments, and process
Pay is offered on an hourly contractor basis with tiered rates by experience. The project requires a quick technical screening and platform assessment before interviews.
Tiered hourly rates published in the project: Junior $9, Middle $12, Senior $16 USD.
Platform pay record lists an average hourly rate of $12.50 USD.
Screening: timed HackerRank assessment plus a platform coding test (candidates should be ready to complete tests within 48 hours of invite).
Interview: recruiter calls and technical follow‑ups for shortlisted candidates.
Who should apply
Apply if you are an experienced Python engineer who enjoys building developer tooling and reliable infrastructure for LLMs and agents. This role fits engineers who like hands‑on coding, automation, mentoring, and close collaboration with researchers.
Good fit: senior engineers who value test‑driven development and infrastructure ownership.
Nice to have: prior exposure to AI model evaluation, scoring pipelines, or agent tooling.
How to get started
Create an OpenTrain profile and apply through the platform so we have your details and can schedule assessments. Be prepared to demonstrate coding and testing skills in the HackerRank and platform tests.
Provide a concise summary of your relevant Python infrastructure experience in your application.
Highlight any LLM/agent evaluation projects, Docker/CI work, and examples of developer tooling you maintain.
Build and own the infrastructure that powers LLM training and agent evaluation at OpenTrain AI; remote for US & Canada, 20+ hours/week, contractor roles with hourly pay tiers (Junior $34 / Middle $37 / Senior $42). Shortlisted candidates complete a timed HackerRank and platform coding test.
Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.