Skip to content
OpenTrain AIFor AI Companies

LLM Code Evaluation Software Engineer

Evaluate and improve AI-generated code while building verification systems for cutting-edge language models. This flexible contractor role is open to software engineers in the US, Canada, and select Western European countries.

OpenTrain AI

Coding & Software

Remote

10 countries

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open to applicants in

United States Canada Austria Belgium France Germany Netherlands Switzerland Luxembourg Ireland

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps people discover opportunities, create a professional profile, and apply to projects that shape how modern AI systems are built.

  • Free account creation
  • Remote opportunities across the AI training industry
  • A place to build experience and grow an AI training career

About AI Code Training

Large language models learn to write better software through carefully prepared examples, expert solutions, and rigorous evaluations. In this growing field, software engineers help identify weaknesses in AI-generated code and create the feedback systems that improve model performance.

  • Work directly on datasets used to train and benchmark AI models
  • Apply software engineering judgment to evaluate model-generated solutions
  • Help shape how AI systems handle real software engineering tasks

The Role

OpenTrain AI is seeking an LLM Code Evaluation Software Engineer to create datasets for training, benchmarking, and advancing large language models. You will work closely with researchers to curate code examples, develop precise solutions, correct code, and assess AI-generated software across multiple programming languages.

This is a contractor, part-time engagement with flexible scheduling. The role requires a minimum commitment of 10 hours per week and may offer up to 40 hours per week. The initial duration is one month, with potential extensions based on performance and fit.

  • Contractor position
  • Flexible schedule
  • Minimum 10 hours per week; engagement details allow up to 40 hours per week
  • Initial one-month term with potential extension
  • No medical or paid leave included

What You'll Do

You will combine hands-on software engineering with structured evaluation of language model output. Your work will support reliable, scalable AI coding solutions and help researchers understand model capabilities throughout the software engineering cycle.

  • Curate code examples and build solutions in Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go.
  • Evaluate and refine AI-generated code for efficiency, scalability, and reliability.
  • Collaborate with cross-functional teams to improve AI coding solutions against industry benchmarks.
  • Build agents that verify code quality and identify recurring error patterns.
  • Form hypotheses about software engineering workflows and evaluate model capabilities at different steps.
  • Design verification mechanisms that automatically validate solutions to software engineering tasks.

Requirements

This role is designed for an experienced software engineer who can assess code deeply and explain technical judgments clearly. Strong full-stack development and production software experience are essential.

  • At least three years of software engineering experience
  • Strong expertise building full-stack applications
  • Experience deploying scalable, production-grade software
  • Deep understanding of software architecture, design, development, and debugging
  • Strong ability to assess code quality and conduct code reviews
  • Excellent oral and written communication skills
  • Ability to provide clear, structured evaluation rationales

Helpful Background

Experience with large language model evaluation or AI model training is helpful, but the core focus is strong software engineering judgment and the ability to evaluate coding solutions.

  • LLM evaluation experience is a plus
  • AI model training experience is a plus
  • Professional experience across multiple programming languages is relevant to the work

Location and Eligibility

Candidates must be based in the United States, Canada, or an eligible Western European country. The role is available in English.

  • United States
  • Canada
  • Austria, Belgium, France, Germany, the Netherlands, Switzerland, Luxembourg, or Ireland
  • English-language work

Why Join AI Training Work

AI training is the human side of building artificial intelligence. Engineers, writers, reviewers, and other specialists provide the examples and feedback that help AI systems become more capable, accurate, and useful.

  • Remote work with flexible hours
  • A chance to contribute to cutting-edge AI development
  • Work that can fit around other professional commitments
  • An opportunity to build experience in a fast-growing technology field

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Senior Software Engineer - C# LLM Evaluation & Code Validation

Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Senior LLM Code Evaluation Engineer

Join OpenTrain AI to build evaluation datasets from public open-source code and measure how LLMs handle real-world software tasks; requires 3+ years software engineering with strong Go skills, 20+ hours/week, and eligibility in select countries.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

C++ LLM Evaluation Software Engineer

Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026