Evaluate frontier AI coding agents on real-world infrastructure tasks—cloud, Kubernetes, CI/CD, observability—and earn $85/hr as a part-time contractor (20+ hrs/wk). Ideal for DevOps/SRE professionals with hands-on cloud and AI-agent experience.
Generative AI & RLHF
Remote Hourly · $85/hr
$85/hr
Compensation
39 countries
Eligibility
Entry
Experience
Jul 29, 2026
Posted
Open to applicants in
Canada United States Costa Rica Guatemala Mexico Panama Denmark Estonia Finland Ireland Latvia Lithuania Norway Sweden Austria Belgium France Germany Netherlands Switzerland United Kingdom Albania Bosnia & Herzegovina Croatia Greece Italy Malta Portugal Serbia Slovenia Spain Bulgaria Czechia Hungary Moldova Poland Romania Slovakia Australia
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover projects, build a unified AI-training portfolio, and grow a durable freelance career focused on the human side of AI.
Creating an OpenTrain account is free. Our projects are remote-friendly and designed for contributors who want flexible, meaningful work that directly shapes how AI systems behave.
About AI training and this work
AI training (also called data labeling or human feedback work) is where human expertise teaches models to perform safely and reliably. Evaluating model outputs is a core part of that process—especially for generative coding agents that propose infrastructure solutions.
This role is part of that human-in-the-loop effort: you will read, test, and judge model-generated infrastructure code and plans so models learn to handle real production systems.
The role
Title: Infrastructure Engineering AI Model Evaluator. You will evaluate frontier code agents and their outputs across realistic infrastructure workflows—cloud platforms, Kubernetes, CI/CD, observability, and automation. Your judgments will focus on bugs, edge cases, reliability issues, and failure modes.
Work type: Contractor, part-time (20+ hours per week).
Labeling task: evaluation/rating of model-generated code and infrastructure solutions.
Data type: computer code / programming; label type: evaluation_rating.
What you'll do
Evaluate frontier AI coding agents on infrastructure engineering tasks and realistic system scenarios.
Review model-generated implementations for cloud, Kubernetes, CI/CD, observability, and infrastructure automation.
Identify bugs, edge cases, reliability issues, security and failure modes in model outputs.
Compare outputs from multiple frontier models and assess relative strengths and weaknesses.
Apply production-minded engineering judgment to rate the safety, correctness, and operational readiness of proposals.
Requirements
You must be able to apply practical infrastructure engineering experience to judge model outputs. We preserve the role's technical demands exactly as listed below.
Minimum 2+ years of professional DevOps, SRE, or cloud engineering experience.
Hands-on experience with one or more: AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling.
Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, or Gemini CLI and comfort evaluating AI-generated code.
Ability to evaluate infrastructure reliability and operational trade-offs in realistic scenarios.
Experience supporting production-scale systems is preferred.
Language: English required.
Eligible countries: CA, US, CR, GT, MX, PA, DK, EE, FI, IE, LV, LT, NO, SE, AT, BE, FR, DE, NL, CH, GB, AL, BA, HR, GR, IT, MT, PT, RS, SI, ES, BG, CZ, HU, MD, PL, RO, SK, AU.
Compensation, schedule, and employment type
This is a contractor role with a paid hourly rate of $85 USD. Expect part-time work at 20+ hours per week and flexible scheduling to fit your availability.
Pay: $85 USD per hour (hourly, contractor).
Expected time commitment: 20+ hours/week (part-time, flexible).
OpenTrain AI seeks senior civil engineers to design scenario-based evaluation tasks and rubrics for AI used in infrastructure and construction; remote contractor role at $70–$80/hr, ~20+ hours/week. Use PE-level judgment across structural, geotech, water resources, transportation, and construction m
Join OpenTrain to evaluate and improve AI-generated civil engineering outputs, providing technical feedback on calculations, assumptions, safety, and professional judgment. Remote contractor role, 20+ hours/week, paid $45–$100/hr; English fluency and 4+ years civil engineering experience required.
Join OpenTrain as a US-licensed civil engineer to review and improve AI-generated civil engineering content, providing expert feedback, references, and practical insight. Flexible, remote contract work paying $65–$90/hr for experienced professionals.