Skip to content
OpenTrain AIFor AI Companies

Software Engineering LLM Evaluator

Use your software engineering judgment to evaluate, correct, and benchmark AI-generated code across the development lifecycle. This flexible contractor role offers 10 to 40 hours weekly with potential extensions.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Jul 16, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. As the hiring and contracting organization for this role, OpenTrain connects skilled contributors with practical projects that help shape how modern AI systems are built.

Create a free OpenTrain account to build a professional profile, showcase relevant experience, discover matching opportunities, and apply in minutes.

About AI Training and Coding Evaluation

AI training is the human side of building artificial intelligence. People prepare examples, review model outputs, and provide structured feedback so AI systems can become more accurate, reliable, and useful.

In this role, your software engineering expertise will directly inform the training and benchmarking of large language models. You will assess whether generated code works well in realistic development contexts and help identify where models need improvement.

The Role

OpenTrain is seeking a Software Engineering LLM Evaluator to create and improve coding datasets used to train and benchmark large language models. The work combines practical software engineering judgment with structured evaluation of model capabilities across the development lifecycle.

You will review AI-generated code, produce precise solutions, correct implementations, and assess software for efficiency, scalability, reliability, maintainability, and alignment with industry performance expectations.

  • Contractor engagement with flexible participation from 10 to 40 hours per week
  • Partial PST overlap is required
  • Initial duration of one month, with potential extensions based on performance and fit
  • Worldwide remote opportunity

What You'll Do

You will curate code examples and build solutions for AI model training initiatives. You will also help create verification systems and collaborate with cross-functional teams to improve AI-driven coding solutions against industry benchmarks.

  • Evaluate and refine AI-generated code for correctness, efficiency, scalability, reliability, and maintainability
  • Work across Python, JavaScript including ReactJS, C/C++, Java, Rust, and Go
  • Build agents that verify code quality and identify recurring error patterns
  • Evaluate model capabilities across prototyping, architecture design, API design, and production implementation
  • Assess software work across launch, experimentation, monitoring, and operational maintenance
  • Design automated verification mechanisms for software engineering tasks
  • Provide structured feedback that improves coding datasets and AI model performance

Requirements

This role requires several years of professional software engineering experience, including at least two continuous years of full-time experience at a top-tier product company. Strong full-stack application development and scalable production deployment experience are required.

Candidates should be comfortable making detailed technical judgments and explaining them clearly in both written and oral communication.

  • Several years of professional software engineering experience
  • At least two continuous years of full-time experience at a top-tier product company
  • Strong full-stack application development experience
  • Experience with scalable production deployment
  • Deep knowledge of software architecture, design, development, debugging, and code quality assessment
  • Clear, structured written and oral evaluation rationales
  • Proficiency in one or more of Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go
  • English-language communication proficiency

Why Do AI Training Work

AI training and data-labeling work is a fast-growing area of technology. Contributors review examples and model behavior that influence the capabilities of state-of-the-art AI systems, making this an opportunity to apply real engineering expertise to emerging tools.

Remote, flexible project structures can make AI training work compatible with other professional commitments while giving specialists a way to build experience in an expanding field.

  • Work remotely with a computer and internet connection
  • Choose a flexible weekly workload within the engagement range
  • Apply software engineering expertise to cutting-edge AI development
  • Build a credible portfolio of AI training and evaluation experience

How to Apply Through OpenTrain

Create or update your free OpenTrain profile with your software engineering background, production experience, technical strengths, and communication skills. Review the opportunity details and apply through OpenTrain in minutes.

Your profile can help you present relevant experience in one place and discover future AI training opportunities aligned with your engineering expertise.

  • Highlight full-stack development and production deployment experience
  • List the programming languages and frameworks you know best
  • Describe your experience with architecture, debugging, and code quality assessment
  • Apply through OpenTrain for consideration

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

LLM Evaluation Software Engineer Ruby

Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.

Coding & Software
Text
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 20, 2026

Senior Python Software Engineer LLM Evaluation

Evaluate how AI models fix real software bugs in open-source Python repositories. This flexible, part-time contractor role focuses on GitHub issue triage, Docker environments, testing, and LLM performance assessment.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

C++ LLM Evaluation Software Engineer

Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026