You will create realistic coding challenges that ask AI systems to implement backend features or fix software bugs. You will also define dependable ways to judge whether generated code works correctly.
The work combines practical backend engineering with the creation of training and evaluation data for AI systems.
- Create feature implementation and bug-fix tasks for backend codebases.
- Design challenging scenarios that test practical engineering judgment.
- Build deterministic verifiers that accept valid solutions and reject invalid ones.
- Assess solution quality and correctness using code review judgment.
- Use debugging, performance analysis, automated testing, and refactoring.
- Work with services, APIs, algorithms, data structures, and existing codebases.
What it pays and takes
This is an individual contractor role with a flexible schedule of approximately 15 hours per week. The listing is marked entry level, while the work requires professional backend development experience and independent technical judgment.
- $30 to $100 per hour.
- Approximately 15 hours per week with a flexible schedule.
- Open to candidates in the eligible countries specified for this role.
- Fluent English.
- Professional backend development experience with Java, Node.js, Go, Rust, TypeScript, Python, C++, or a similar language.
- Experience building scalable services and APIs.
- Ability to diagnose bugs and system bottlenecks.
- Experience implementing complex features and refactoring large-scale applications.
- Strong knowledge of algorithms, data structures, and foundational computer science.
- Comfort with automated testing, code review, analytical problem-solving, and technical communication.
- Experience with unit tests, test strategies, flaky tests, or continuous integration test pipelines is helpful.
- Prior AI training experience is not required.
How it works
Apply on OpenTrain with your resume, then complete the application on the hiring site.
About AI training work
AI training work is the human work behind modern AI systems. People create examples, write evaluations, and review model outputs so systems can produce more useful and reliable results.