Skip to content
OpenTrain AIFor AI Companies

Python Engineer, AI Coding-Tool Evaluation (Part-Time)

Part-time contractor role testing and evaluating an internal AI coding tool; $100/hr, remote, under 20 hrs/week. Ideal for intermediate Python engineers with hands-on experience using AI coding assistants like Cursor, Windsurf, or Claude Code.

OpenTrain AI

Coding & Software

100% Remote Hourly · $100/hr

$100/hr

Compensation

Worldwide

Eligibility

Intermediate

Experience

Oct 17, 2025

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for people who start and grow careers in AI training and data labeling. We connect skilled contributors with paid, remote projects that help shape how modern AI systems behave.

Work with OpenTrain to do cutting-edge annotation and evaluation tasks on flexible schedules — ideal for part-time contributors, developers, and domain specialists.

About This Project

This project supports an AI research effort backed by $10M in funding. The team includes professors, serial entrepreneurs, and AI researchers from top institutions and industry research groups.

You will test and evaluate an internal coding tool that helps developers write and review code. Your feedback will directly improve the tool's accuracy, UX, and code-generation behavior.

The Role

OpenTrain is hiring an intermediate-level Python engineer to join as a part-time contractor and evaluate an internal AI coding tool. This is a hands-on testing and annotation role where your technical judgment matters.

Work is performed using the project's proprietary tooling to label and rate code outputs, reproduce issues, and provide structured feedback.

  • Commitment: Less than 20 hours per week (flexible)
  • Pay: $100 USD per hour (contractor)
  • Employment type: Contractor, Part-time
  • Location: Remote — worldwide applicants welcome
  • Data type: Computer code / programming
  • Label types: Code authoring/annotation and evaluation/rating
  • Tooling: Internal proprietary annotation/testing tooling

What You'll Do

Your day-to-day work focuses on exercising the coding tool, producing and reviewing code samples, and rating tool outputs for correctness, style, and usefulness.

  • Use the internal tool to author, edit, and evaluate code snippets and programmatic solutions.
  • Rate and annotate generated code according to project guidelines (accuracy, security, readability).
  • Reproduce bugs and edge cases, submit clear issue reports, and suggest improvements.
  • Perform comparative evaluations across different prompts or tool behaviors.
  • Provide structured feedback on UX, developer workflow fit, and failure modes.

Requirements

Candidates must meet the technical and practical requirements below to be considered.

  • Solid Python experience (back-end or full-stack development experience required)
  • Hands-on familiarity with AI coding assistants or coding tools such as Cursor, Windsurf, and Claude Code
  • Intermediate experience level: able to evaluate code correctness, debug, and give technical feedback
  • Reliable internet connection and ability to work remotely using provided proprietary tooling
  • Comfort working as a contractor and logging hours for hourly payment

Who Should Apply

This role is a fit for engineers who enjoy both coding and critically evaluating developer-facing AI tools. You should be comfortable reading, writing, and assessing code across common Python use-cases and communicating concise, actionable feedback.

  • Back-end or full-stack Python engineers with experience using AI coding assistants
  • Engineers who like exploratory testing, reproducing edge cases, and improving tooling
  • People seeking flexible, part-time contract work in the AI-training field

How It Works / Next Steps

Apply through OpenTrain to be considered. If selected, you'll complete a short onboarding and qualification task to confirm skill fit and access the proprietary testing environment.

As a contractor you will log hours and be paid $100/hr. The project runs with flexible scheduling under 20 hours per week.

  • Application -> qualification task -> onboarding -> paid contractor work
  • You will use internal proprietary tooling for labeling, evaluation, and reporting
  • OpenTrain manages contracting and payments for this role

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all jobs

Python Backend Developer – AI Coding Evaluation

Build scalable Python APIs and help test AI coding tools in focused four-day bursts. This remote contractor role offers $50–$100 per hour, 20+ hours weekly, and the chance to shape developer-focused AI systems.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $50–$100/hr

Posted Jun 28, 2026

Senior Python Developer for AI Model Evaluation

Use advanced Python skills to evaluate AI models, create training data, rank responses, and improve coding-focused systems. This worldwide contract role offers flexible part-time work of 20+ hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026

AI Coding Agent Evaluation Expert

Help improve AI coding assistants by evaluating complex technical work, from large-codebase debugging to architectural changes and end-to-end features. This worldwide, part-time contractor role offers $80–$100 per hour and requires 20+ hours weekly.

Coding & Software
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $80–$100/hr

Posted Aug 7, 2026