Skip to content
OpenTrain AIFor AI Companies

LLM Code Evaluation Software Engineer

Help train and benchmark large language models by evaluating, correcting, and verifying code across Python, JavaScript, C/C++, Java, Rust, and Go. This flexible contractor role is open to software engineers in the US, Canada, and Western Europe.

OpenTrain AI

Coding & Software

Remote

6 countries

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open to applicants in

United States Canada Austria
+3 more
  • Austria
  • Belgium
  • Canada
  • France
  • Germany
  • United States

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover projects, build a professional profile, and apply to opportunities that shape the future of artificial intelligence. Creating an OpenTrain account is free.

About AI Training Work

AI training is the human side of building modern artificial intelligence. Software engineers contribute by creating examples, reviewing model-generated code, rating results, and developing reliable ways to test whether AI systems can solve real engineering problems.

This work offers a direct opportunity to influence cutting-edge language models while working remotely and flexibly. Contributors with specialized technical expertise can take on projects that require advanced knowledge of software development and code quality.

  • Help improve how large language models understand and generate software.
  • Work remotely with a flexible schedule within the engagement requirements.
  • Apply your engineering judgment to training, benchmarking, and evaluation data.

The Role

OpenTrain is hiring a contractor LLM Code Evaluation Software Engineer to help create datasets for training, benchmarking, and advancing large language models. You will work with researchers and cross-functional teams to curate code examples, develop precise solutions, correct AI-generated code, and evaluate model capabilities across software engineering tasks.

The role focuses on producing dependable coding solutions and designing verification systems that can identify errors and assess quality. Prior experience with LLM evaluation or AI model training is helpful but not required.

  • Contractor, part-time engagement
  • Flexible schedule with a stated minimum of 10 hours and up to 40 hours per week; structured details indicate 20+ hours per week
  • Initial duration of 1 month, with potential extensions based on performance and fit
  • Candidates must be based in the US, Canada, or listed Western European countries
  • English-language work

What You'll Do

  • Curate code examples and build solutions in Python, JavaScript including ReactJS, C/C++, Java, Rust, and Go.
  • Evaluate and refine AI-generated code for efficiency, scalability, and reliability.
  • Collaborate with cross-functional teams to improve AI coding solutions against industry benchmarks.
  • Build agents that verify code quality and identify recurring error patterns.
  • Form hypotheses about steps in the software engineering cycle and evaluate model capabilities at those steps.
  • Design verification mechanisms that automatically assess solutions to software engineering tasks.
  • Provide precise, clear, and structured rationales for evaluation decisions.

Requirements

This role requires several years of software engineering experience, with at least 3 years indicated in the requirements. You should be comfortable building and deploying production-grade software and explaining technical judgments clearly in both writing and conversation.

  • At least 3 years of software engineering experience.
  • Strong expertise building full-stack applications.
  • Experience deploying scalable, production-grade software.
  • Deep understanding of software architecture, design, development, and debugging.
  • Strong ability to assess code quality and conduct or interpret code reviews.
  • Excellent oral and written communication skills for clear evaluation rationales.

Who Should Apply

This opportunity is designed for software engineers who want to apply their development expertise to AI training and evaluation. It may be a strong fit if you enjoy analyzing how systems work, identifying error patterns, designing automated checks, and translating technical reasoning into consistent evaluation criteria.

  • Software engineers with full-stack and production deployment experience.
  • Developers who work with one or more of Python, JavaScript, ReactJS, C/C++, Java, Rust, or Go.
  • Engineers interested in benchmarking and improving large language models.
  • Candidates with previous LLM evaluation or AI model training experience, which is a plus.
  • Applicants based in the United States, Canada, Austria, Belgium, France, Germany, or other eligible Western European countries.

Engagement Details

This is a flexible, part-time contractor engagement. The initial term is one month, with possible extensions based on performance and fit. The role does not include medical or paid leave.

  • Employment type: Contractor and part-time
  • Location: Remote, limited to eligible countries
  • Duration: 1 month initially, with potential extension
  • Schedule: Flexible, up to 40 hours per week
  • Compensation: USD terms are listed, but no rate is specified in the provided details
  • Benefits: No medical or paid leave

How to Apply Through OpenTrain

Create a free OpenTrain account, build your profile around your software engineering experience, and apply in minutes. OpenTrain helps contributors start and grow careers in the fast-moving AI training and data-labeling industry, where human technical judgment remains essential to building capable and reliable AI systems.

  • Highlight your full-stack development and production deployment experience.
  • List the programming languages and frameworks you use confidently.
  • Describe your experience with debugging, code review, architecture, or automated verification.
  • Mention LLM evaluation or AI training experience if applicable.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

C# Software Engineer LLM Code Evaluation

Evaluate how language models solve real software engineering problems using C#, GitHub repositories, Docker, and unit-test analysis. This remote contractor role requires 20 or more hours weekly and four hours of daily PST overlap.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

C++ LLM Evaluation Software Engineer

Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

LLM Code Evaluation Engineer

Evaluate how large language models solve realistic coding and bug-fixing tasks across open-source repositories. This worldwide, part-time contractor role offers hands-on AI training work for engineers comfortable with Git, Docker, and real codebases.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026