Build Python infrastructure for LLM training and agent evaluation, including secure sandboxes, task frameworks, scoring pipelines, and developer tools. This remote contract role offers 20+ hours per week for engineers in the United States or Canada.
The work
You will build and maintain the infrastructure behind LLM training and agent evaluation. Your work will help AI researchers create tasks, run tests, and assess how well agents perform.
The role combines Python engineering with developer tooling, automated evaluation, and secure execution environments.
- Build secure sandboxes for running agent tasks.
- Create reusable task frameworks and scoring pipelines.
- Develop automated evaluation pipelines for coding and agent tasks.
- Build reusable repositories and modern developer environments.
- Use Python back ends with FastAPI or Flask where needed.
- Guide AI researchers in using the tools you build.
What it pays and takes
This is a part-time contractor role for engineers who can take ownership of infrastructure work. Compensation is tiered by experience.
Clear English communication and a collaborative working style are required. Previous work on coding-task platforms is a strong bonus.
- Pay: Junior $34 per hour, Middle $37 per hour, or Senior $42 per hour in USD.
- Time: 20 or more hours per week.
- Location: Fully remote for people located in the United States or Canada.
- Language: Clear English communication.
- Experience: 5 or more years of hands-on Python experience.
- Required skills: Deep Linux and Docker experience.
- Required skills: Designing CI/CD pipelines, preferably with GitHub Actions.
- Required skills: FastAPI or Flask back ends, pytest-based testing, devcontainers, and Makefiles.
- Experience level: Intermediate, with a senior-minded approach to ownership and technical guidance.
- Selection requirement: Short-listed candidates complete a timed HackerRank assessment and a platform coding test before recruiter interviews.
How it works
Apply on OpenTrain. The employer reviews applications there.
About AI training work
AI training is the human work used to improve AI systems, including writing coding tasks, testing model behavior, and reviewing agent results. Engineers with strong software skills are needed to build the tools and evaluation systems that make this work reliable.