Evaluate and improve AI-generated code while creating benchmarking datasets for advanced software engineering models. This expert, remote contract role is part time at less than 20 hours per week.
Coding & Software
100% Remote
Worldwide
Eligibility
Expert
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps people discover projects, build a professional profile, and apply to opportunities that advance their work in this fast-growing field.
Free OpenTrain account
Remote opportunity open worldwide
Part-time contractor engagement
About AI Training and Code Evaluation
AI training is the human side of building artificial intelligence. Software experts help modern models improve by preparing coding examples, reviewing generated solutions, and explaining why code is effective, scalable, reliable, or incorrect.
This work supports the development of AI systems that assist with software engineering. Your evaluations and verification methods can help shape how models generate, debug, and assess production-quality code.
Work directly on cutting-edge AI training and benchmarking
Apply advanced software engineering judgment to model outputs
Contribute remotely with a flexible part-time schedule
The Software Engineering Code Evaluator Role
OpenTrain is hiring an expert Software Engineering Code Evaluator to create training and benchmarking datasets for large language models. The role focuses on curating code examples, writing precise solutions, correcting code, and evaluating AI-generated software across multiple programming languages and frameworks.
You will also help build verification agents, identify recurring error patterns, and design mechanisms that can automatically verify solutions to software engineering tasks.
Experience level: Expert
Less than 20 hours per week
Contractor, part-time position
English-language work
Open worldwide
What You'll Do
You will combine hands-on software engineering expertise with structured evaluation to help create reliable coding datasets and benchmarks. The work includes both detailed code review and the design of systems that assess coding solutions automatically.
Curate code examples and produce corrected solutions for software engineering tasks
Review AI-generated code for efficiency, scalability, and reliability
Build agents that assess code quality and detect recurring error patterns
Design automatic verification mechanisms for coding solutions
Collaborate with researchers and cross-functional teams on model benchmarking and coding evaluations
Write clear, structured rationales for code review and evaluation decisions
Required Experience and Skills
This role is intended for an experienced software engineer with a strong background in production software development, full-stack applications, debugging, architecture, and code review. You should be comfortable making precise technical judgments and communicating the reasoning behind them.
Several years of software engineering experience
At least 2 years of continuous full-time experience at a top-tier product company
Strong experience building full-stack applications and production-grade software with modern languages and tools
Deep understanding of software architecture, development, debugging, and code review standards
Experience curating or correcting programming solutions for model training or evaluation
Strong ability to review code for efficiency, scalability, and reliability
Familiarity with Python, JavaScript, ReactJS, C/C++, Java, Rust, or Go
Experience designing verification logic or automated tests for software tasks
Excellent oral and written communication skills
Clear written rationales for code review and evaluation decisions
Why This Work Matters
Every major AI system depends on examples and expert feedback prepared by people. By evaluating code and developing reliable verification methods, you will contribute to how state-of-the-art models behave in software development settings.
Use your engineering expertise in an emerging AI field
Help improve the quality and reliability of coding models
Build experience in AI training, benchmarking, and model evaluation
How to Get Started
Create a free OpenTrain account to build your profile and apply in minutes. Your software engineering background, programming experience, and ability to explain technical evaluations will be central to your application.
Prepare details about your software engineering experience
Highlight full-stack and production software work
Showcase experience with code review, automated testing, or verification logic
Include relevant programming languages and frameworks
Evaluate and improve AI-generated code while building verification systems for cutting-edge language models. This flexible contractor role is open to software engineers in the US, Canada, and select Western European countries.
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.