Review AI-generated code, coding-agent behavior, and tool use across real software repositories. This remote contractor role seeks experienced engineers for at least 20 hours per week, with six hours of Pacific Time overlap preferred.
The work
You will assess AI-generated code and coding-agent behavior in substantial software repositories. Your reviews will help improve coding models by turning technical judgment into clear evaluation signals and feedback.
You will work with researchers and engineers to make evaluation processes more consistent and scalable.
- Review generated solutions, code changes, and tool use for correctness, robustness, and maintainability.
- Compare model outputs and explain why one solution is stronger than another.
- Find technical errors, weak approaches, and recurring model failure patterns.
- Create and improve rubrics and evaluation criteria for coding tasks.
- Produce evaluation and preference data for coding-model improvement.
- Support data-generation, collection, and evaluation pipelines and infrastructure.
- Write clear findings, recommendations, and practical updates.
What it pays and takes
This is a remote independent contractor engagement expected to last about three months. Compensation is market rate, with hourly pay based on the rate agreed for the engagement.
- Pay: Market-rate hourly compensation at an agreed rate; no specific rate is provided.
- Hours: At least 20 hours per week; a 40-hour week is preferred.
- Time zone: At least six hours of Pacific Time overlap is preferred.
- Location: The description lists eligible professionals in North America, LATAM, or India. The listing metadata specifies India.
- Contract: Independent contractor and part-time engagement.
- Experience: At least five years of hands-on software engineering experience is required, although the listing is tagged entry level.
- Technical skills: Strong Python, TypeScript, JavaScript, Go, or another major production language, plus experience with substantial real-world codebases.
- Judgment and communication: Strong code-review skills, precise technical judgment, and the ability to explain whether an implementation is correct or how it could improve.
- AI tools: Experience using modern large language models or AI coding tools is required.
- Helpful background: Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is useful but not required.
- Language: English.
How it works
Apply on OpenTrain with your resume, then complete the application on the hiring site.
About AI training work
AI training work uses human reviews, examples, and feedback to improve artificial intelligence systems. In this role, your software engineering judgment helps models produce more correct, robust, and maintainable code, which is why hands-on experience matters.
- OpenTrain is the hiring and contracting organization for this role and supports careers in AI training and data labeling.
- AI training can include reviewing model outputs, rating alternatives, writing feedback, and creating evaluation data.