Review AI-generated code for correctness, security, reliability, and maintainability in LLM evaluation work. This remote contract role requires at least seven years of software engineering experience and 20+ hours per week.
The Work
You will review AI-generated code used to train and evaluate large language models. Your work will cover different programming languages and software situations, including bug fixes, new features, refactoring, API integrations, configuration changes, and database operations.
You will assess whether code works as intended and identify defects, logic errors, missing pieces, edge-case failures, performance problems, security risks, and architecture weaknesses. You will compare possible implementations, improve code into reliable reference solutions, explain your recommendations, and help create technical rubrics and coding benchmarks.
- Review AI-generated code and software changes for correctness, security, reliability, scalability, readability, and maintainability.
- Debug complex codebases and identify the root causes of technical problems.
- Compare implementations and select or create the solution that best meets the technical requirements.
- Write clear technical feedback and contribute to evaluation criteria, coding benchmarks, and improved LLM evaluation methods.
What It Pays And Takes
The role details do not list a pay rate. This is a remote, part-time contract role for an individual contributor supporting AI training and engineering evaluation work.
- Pay: Not provided in the role details.
- Time: 20+ hours per week.
- Location: Worldwide and remote.
- Language: Written English proficiency is required.
- Experience: The listing is marked entry level in the source fields, while the role description requires at least seven years of professional software engineering experience.
- Programming: Strong proficiency in at least one of Python, JavaScript, TypeScript, Java, C++, Go, C#, Ruby, PHP, or Rust.
- Technical knowledge: Production software development, debugging, code review, clean code, modular architecture, abstraction, error handling, data structures, algorithms, APIs, databases, and application architecture.
- Practices: Experience with collaborative code reviews, Git, and modern software engineering methods.
- Helpful background: Experience evaluating AI-generated code, creating technical rubrics, contributing to software engineering benchmarks, or working across multiple languages and architecture patterns.
How It Works
Apply on OpenTrain with your resume and then complete the application on the hiring site.
About AI Training Work
AI training is the human work behind systems that generate and understand code, text, images, and other data. People review examples, rate model outputs, and provide clear corrections so AI systems become more useful and reliable; OpenTrain helps people find and build careers in this field.