Skip to content
OpenTrain AIFor AI Companies

Senior Software Engineer LLM Evaluation

Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free, and this opportunity gives experienced software engineers a way to contribute directly to the development of advanced AI systems.

  • Contractor position
  • Part-time engagement
  • Worldwide opportunity
  • English-language work
  • 20+ hours per week

About AI Training and LLM Evaluation

AI training is the human side of building artificial intelligence. People prepare examples, write and review responses, and evaluate model outputs so that modern AI systems become more accurate, reliable, and useful.

In this role, your software engineering expertise will help train and benchmark large language models that generate code. You will assess their capabilities across software engineering tasks and help create verification methods that measure quality, efficiency, and reliability.

  • Work on cutting-edge large language model development
  • Shape how AI systems generate and evaluate software
  • Contribute remotely with flexible part-time hours

The Role

OpenTrain AI is recruiting a Senior Software Engineer for LLM Evaluation. In this contractor role, you will create and refine high-quality datasets for training and benchmarking large language models.

You will curate code examples, build and correct solutions in multiple programming languages, and evaluate AI-generated code for efficiency and reliability. You will also work with cross-functional teams to improve AI-driven coding solutions against industry performance benchmarks.

  • Subject matter: LLM code evaluation and dataset curation
  • Primary focus: software engineering evaluation and programming data
  • Listed experience level: entry level, with substantial professional experience required in the qualifications

What You’ll Do

Your work will combine hands-on software development with structured evaluation of AI-generated solutions. You will need to explain your judgments clearly and design processes that can assess software engineering performance consistently.

  • Curate code examples for AI model training initiatives.
  • Build and correct solutions in Python, JavaScript including ReactJS, C/C++, Java, Rust, and Go.
  • Evaluate and refine AI-generated code for efficiency, scalability, and reliability.
  • Collaborate with cross-functional teams to improve AI-driven coding solutions against industry performance benchmarks.
  • Build agents that verify code quality and identify error patterns.
  • Hypothesize about steps in the software engineering cycle and evaluate model capabilities at each step.
  • Design verification mechanisms that can automatically verify solutions to software engineering tasks.
  • Write clear, structured rationales explaining evaluation decisions.

Requirements

This role requires several years of software engineering experience, including at least two years of continuous full-time experience at a top-tier product company. The listed examples include Google, Stripe, Amazon, Apple, Meta, Netflix, and Microsoft.

The role also calls for strong full-stack development experience and the ability to deploy scalable, production-grade software using modern languages and tools.

  • Several years of software engineering experience.
  • At least 2 years of continuous full-time experience at a top-tier product company such as Google, Stripe, Amazon, Apple, Meta, Netflix, or Microsoft.
  • Strong expertise building full-stack applications.
  • Experience deploying scalable, production-grade software.
  • Deep understanding of software architecture, design, development, debugging, and code quality review.
  • Excellent oral and written communication skills.
  • Ability to write clear, structured evaluation rationales.
  • Experience with Python, JavaScript, ReactJS, C/C++, Java, Rust, and/or Go is essential for success.

Who Should Apply

This opportunity is suited to software engineers who can combine deep technical judgment with careful written evaluation. It may be especially relevant to engineers who have built full-stack systems, reviewed production code, debugged complex software, or worked with several of the listed programming languages.

You should be comfortable assessing not only whether code works, but also whether it is scalable, efficient, reliable, and aligned with sound software engineering practices.

  • Experienced full-stack software engineers
  • Engineers with strong architecture and code review backgrounds
  • Developers comfortable evaluating code across multiple languages
  • Clear technical communicators who can justify structured judgments
  • Candidates able to commit to at least 20 hours per week

How the Engagement Works

This is a worldwide, part-time contractor opportunity conducted in English. The time requirement is 20 or more hours per week, making it suitable for professionals seeking flexible remote work while contributing to the development of advanced AI coding systems.

Apply through OpenTrain to create your profile and submit your application. OpenTrain accounts are free, and the platform supports contributors as they build careers in AI training and data labeling.

  • Remote and open worldwide
  • Part-time contractor engagement
  • 20+ hours per week
  • English required
  • Apply through OpenTrain

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Senior Software Engineer, LLM Evaluation

Build realistic coding-agent benchmarks, test suites, and security-focused evaluators for large language models. This part-time contractor role is open worldwide and requires senior production engineering experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Aug 6, 2026

Senior Python Software Engineer LLM Evaluation

Evaluate how AI models fix real software bugs in open-source Python repositories. This flexible, part-time contractor role focuses on GitHub issue triage, Docker environments, testing, and LLM performance assessment.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

Senior C++ Software Engineer for LLM Evaluation

Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026