Skip to content
OpenTrain AIFor AI Companies

AI Coding Model Evaluation Engineer

Review AI-generated code, coding-agent behavior, and tool use across real software repositories. This remote contractor role seeks experienced engineers for at least 20 hours per week, with six hours of Pacific Time overlap preferred.

Apply now
OpenTrain AI

Coding & Software

Remote

1 country

Eligibility

Entry

Experience

Sep 24, 2026

Posted

Open to applicants in

India

The work

You will assess AI-generated code and coding-agent behavior in substantial software repositories. Your reviews will help improve coding models by turning technical judgment into clear evaluation signals and feedback.

You will work with researchers and engineers to make evaluation processes more consistent and scalable.

  • Review generated solutions, code changes, and tool use for correctness, robustness, and maintainability.
  • Compare model outputs and explain why one solution is stronger than another.
  • Find technical errors, weak approaches, and recurring model failure patterns.
  • Create and improve rubrics and evaluation criteria for coding tasks.
  • Produce evaluation and preference data for coding-model improvement.
  • Support data-generation, collection, and evaluation pipelines and infrastructure.
  • Write clear findings, recommendations, and practical updates.

What it pays and takes

This is a remote independent contractor engagement expected to last about three months. Compensation is market rate, with hourly pay based on the rate agreed for the engagement.

  • Pay: Market-rate hourly compensation at an agreed rate; no specific rate is provided.
  • Hours: At least 20 hours per week; a 40-hour week is preferred.
  • Time zone: At least six hours of Pacific Time overlap is preferred.
  • Location: The description lists eligible professionals in North America, LATAM, or India. The listing metadata specifies India.
  • Contract: Independent contractor and part-time engagement.
  • Experience: At least five years of hands-on software engineering experience is required, although the listing is tagged entry level.
  • Technical skills: Strong Python, TypeScript, JavaScript, Go, or another major production language, plus experience with substantial real-world codebases.
  • Judgment and communication: Strong code-review skills, precise technical judgment, and the ability to explain whether an implementation is correct or how it could improve.
  • AI tools: Experience using modern large language models or AI coding tools is required.
  • Helpful background: Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is useful but not required.
  • Language: English.

How it works

Apply on OpenTrain with your resume, then complete the application on the hiring site.

About AI training work

AI training work uses human reviews, examples, and feedback to improve artificial intelligence systems. In this role, your software engineering judgment helps models produce more correct, robust, and maintainable code, which is why hands-on experience matters.

  • OpenTrain is the hiring and contracting organization for this role and supports careers in AI training and data labeling.
  • AI training can include reviewing model outputs, rating alternatives, writing feedback, and creating evaluation data.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Code Generation And Model Evaluation Research Engineer

Review and improve code-generation systems across Python, Java, Rust, C++, Go, and TypeScript. Work remotely for 20+ hours each week on coding tasks, code review, and AI model evaluation.

Coding & Software
Computer Code Programming
Remote · Andorra, United Arab Emirates, Antigua & Barbuda +227 more
English
Part-time · Flexible
Entry level
Hourly · $50–$100/hr

Posted Sep 14, 2026

Software Engineering LLM Evaluator

Review and improve AI-generated code across major programming languages, from design through production operations. This flexible contractor role requires strong software engineering experience and partial PST availability.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

AI Evaluation Engineer, Engineering Simulation

Build and validate demanding engineering simulations that test AI agents across electrical, mechanical, control systems, aerospace, systems, and robotics applications. This remote contractor role requires strong Python, simulation, and engineering design experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 11, 2026