Python Engineer, AI Coding-Tool Evaluation (Part-Time)
Part-time contractor role testing and evaluating an internal AI coding tool; $100/hr, remote, under 20 hrs/week. Ideal for intermediate Python engineers with hands-on experience using AI coding assistants like Cursor, Windsurf, or Claude Code.
Coding & Software
100% Remote Hourly · $100/hr
$100/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Oct 17, 2025
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people who start and grow careers in AI training and data labeling. We connect skilled contributors with paid, remote projects that help shape how modern AI systems behave.
Work with OpenTrain to do cutting-edge annotation and evaluation tasks on flexible schedules — ideal for part-time contributors, developers, and domain specialists.
About This Project
This project supports an AI research effort backed by $10M in funding. The team includes professors, serial entrepreneurs, and AI researchers from top institutions and industry research groups.
You will test and evaluate an internal coding tool that helps developers write and review code. Your feedback will directly improve the tool's accuracy, UX, and code-generation behavior.
The Role
OpenTrain is hiring an intermediate-level Python engineer to join as a part-time contractor and evaluate an internal AI coding tool. This is a hands-on testing and annotation role where your technical judgment matters.
Work is performed using the project's proprietary tooling to label and rate code outputs, reproduce issues, and provide structured feedback.
Commitment: Less than 20 hours per week (flexible)
Pay: $100 USD per hour (contractor)
Employment type: Contractor, Part-time
Location: Remote — worldwide applicants welcome
Data type: Computer code / programming
Label types: Code authoring/annotation and evaluation/rating
Your day-to-day work focuses on exercising the coding tool, producing and reviewing code samples, and rating tool outputs for correctness, style, and usefulness.
Use the internal tool to author, edit, and evaluate code snippets and programmatic solutions.
Rate and annotate generated code according to project guidelines (accuracy, security, readability).
Reproduce bugs and edge cases, submit clear issue reports, and suggest improvements.
Perform comparative evaluations across different prompts or tool behaviors.
Provide structured feedback on UX, developer workflow fit, and failure modes.
Requirements
Candidates must meet the technical and practical requirements below to be considered.
Solid Python experience (back-end or full-stack development experience required)
Hands-on familiarity with AI coding assistants or coding tools such as Cursor, Windsurf, and Claude Code
Intermediate experience level: able to evaluate code correctness, debug, and give technical feedback
Reliable internet connection and ability to work remotely using provided proprietary tooling
Comfort working as a contractor and logging hours for hourly payment
Who Should Apply
This role is a fit for engineers who enjoy both coding and critically evaluating developer-facing AI tools. You should be comfortable reading, writing, and assessing code across common Python use-cases and communicating concise, actionable feedback.
Back-end or full-stack Python engineers with experience using AI coding assistants
Engineers who like exploratory testing, reproducing edge cases, and improving tooling
People seeking flexible, part-time contract work in the AI-training field
How It Works / Next Steps
Apply through OpenTrain to be considered. If selected, you'll complete a short onboarding and qualification task to confirm skill fit and access the proprietary testing environment.
As a contractor you will log hours and be paid $100/hr. The project runs with flexible scheduling under 20 hours per week.
Application -> qualification task -> onboarding -> paid contractor work
You will use internal proprietary tooling for labeling, evaluation, and reporting
OpenTrain manages contracting and payments for this role
Join OpenTrain AI as a part-time contractor building reinforcement-learning environments and reproducible software tasks that evaluate model programming ability — remote, worldwide, under 20 hrs/week, paying up to $150/hr.
Join OpenTrain AI as a Python Data Trainer to create and debug Python code that teaches models; 20–40 hours/week, 3–6 month contract at $8/hr. Ideal for Python developers with 2+ years' experience and strong English.