Skip to content
OpenTrain AIFor AI Companies

Python Engineer, AI Coding-Tool Evaluation (Part-Time)

Part-time contractor role testing and evaluating an internal AI coding tool; $100/hr, remote, under 20 hrs/week. Ideal for intermediate Python engineers with hands-on experience using AI coding assistants like Cursor, Windsurf, or Claude Code.

OpenTrain AI

Coding & Software

100% Remote Hourly · $100/hr

$100/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Oct 17, 2025

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for people who start and grow careers in AI training and data labeling. We connect skilled contributors with paid, remote projects that help shape how modern AI systems behave.

Work with OpenTrain to do cutting-edge annotation and evaluation tasks on flexible schedules — ideal for part-time contributors, developers, and domain specialists.

About This Project

This project supports an AI research effort backed by $10M in funding. The team includes professors, serial entrepreneurs, and AI researchers from top institutions and industry research groups.

You will test and evaluate an internal coding tool that helps developers write and review code. Your feedback will directly improve the tool's accuracy, UX, and code-generation behavior.

The Role

OpenTrain is hiring an intermediate-level Python engineer to join as a part-time contractor and evaluate an internal AI coding tool. This is a hands-on testing and annotation role where your technical judgment matters.

Work is performed using the project's proprietary tooling to label and rate code outputs, reproduce issues, and provide structured feedback.

  • Commitment: Less than 20 hours per week (flexible)
  • Pay: $100 USD per hour (contractor)
  • Employment type: Contractor, Part-time
  • Location: Remote — worldwide applicants welcome
  • Data type: Computer code / programming
  • Label types: Code authoring/annotation and evaluation/rating
  • Tooling: Internal proprietary annotation/testing tooling

What You'll Do

Your day-to-day work focuses on exercising the coding tool, producing and reviewing code samples, and rating tool outputs for correctness, style, and usefulness.

  • Use the internal tool to author, edit, and evaluate code snippets and programmatic solutions.
  • Rate and annotate generated code according to project guidelines (accuracy, security, readability).
  • Reproduce bugs and edge cases, submit clear issue reports, and suggest improvements.
  • Perform comparative evaluations across different prompts or tool behaviors.
  • Provide structured feedback on UX, developer workflow fit, and failure modes.

Requirements

Candidates must meet the technical and practical requirements below to be considered.

  • Solid Python experience (back-end or full-stack development experience required)
  • Hands-on familiarity with AI coding assistants or coding tools such as Cursor, Windsurf, and Claude Code
  • Intermediate experience level: able to evaluate code correctness, debug, and give technical feedback
  • Reliable internet connection and ability to work remotely using provided proprietary tooling
  • Comfort working as a contractor and logging hours for hourly payment

Who Should Apply

This role is a fit for engineers who enjoy both coding and critically evaluating developer-facing AI tools. You should be comfortable reading, writing, and assessing code across common Python use-cases and communicating concise, actionable feedback.

  • Back-end or full-stack Python engineers with experience using AI coding assistants
  • Engineers who like exploratory testing, reproducing edge cases, and improving tooling
  • People seeking flexible, part-time contract work in the AI-training field

How It Works / Next Steps

Apply through OpenTrain to be considered. If selected, you'll complete a short onboarding and qualification task to confirm skill fit and access the proprietary testing environment.

As a contractor you will log hours and be paid $100/hr. The project runs with flexible scheduling under 20 hours per week.

  • Application -> qualification task -> onboarding -> paid contractor work
  • You will use internal proprietary tooling for labeling, evaluation, and reporting
  • OpenTrain manages contracting and payments for this role

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Code Evaluation and Benchmarking Engineer

Evaluate and benchmark AI-generated code: review correctness, debug and verify solutions, and build evaluation datasets for frontier models. US-remote, contractor role — 20+ hrs/week (min 4 hrs/day), 1-month contract with 4-hour PST overlap required.

Coding & Software
Text
Remote · United States
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

Software Engineer for AI Training Environments

Join OpenTrain AI as a part-time contractor building reinforcement-learning environments and reproducible software tasks that evaluate model programming ability — remote, worldwide, under 20 hrs/week, paying up to $150/hr.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $50–$150/hr

Posted Jul 23, 2026

Python AI Data Trainer, Coding & Debugging (20–40 hrs/wk)

Join OpenTrain AI as a Python Data Trainer to create and debug Python code that teaches models; 20–40 hours/week, 3–6 month contract at $8/hr. Ideal for Python developers with 2+ years' experience and strong English.

Coding & Software
Computer Code Programming
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $8/hr

Posted Jun 18, 2024