Skip to content
OpenTrain AIFor AI Companies

Software Engineering AI Evaluation Specialist

Evaluate AI-powered developer workflows through hands-on coding environments, GitHub, CI/CD, and technical assessment. This worldwide contractor role pays $50-$70 per hour for 20+ hours weekly.

OpenTrain AI

Coding & Software

100% Remote Hourly · $50–$70/hr

$50–$70/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Aug 19, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps contributors discover specialized projects, build a profile, and grow a durable portfolio in a rapidly expanding field.

Creating an OpenTrain account is free, and candidates can apply for opportunities that match their technical experience.

  • Worldwide opportunity with remote work
  • Contractor and part-time engagement
  • 20+ hours per week
  • USD $50-$70 per hour

About AI Training Work

AI training is the human side of building artificial intelligence. Contributors prepare examples, review model behavior, and provide structured feedback so AI systems can reason more accurately and perform useful tasks.

In this role, your software engineering expertise will help evaluate developer-facing AI workflows, verify technical outputs, and improve how next-generation models work in realistic development environments.

  • Work directly with cutting-edge AI systems
  • Apply professional engineering judgment to model evaluation
  • Help improve AI behavior through consistent technical feedback

The Role

OpenTrain is seeking a Software Engineering AI Evaluation Specialist to support realistic developer activities and structured technical assessment. You will evaluate AI-powered workflows involving software development, verify model outputs in hands-on environments, and provide consistent judgments that improve model performance.

The work combines professional software engineering with reproducible test-environment design, service integration troubleshooting, and disciplined evaluation of developer-facing AI behavior.

  • Contractor role
  • Part-time work requiring 20+ hours per week
  • English-language work
  • Worldwide eligibility

What You’ll Do

You will execute evaluation assignments across common software development workflows and assess technical results directly in relevant environments. You will also create reliable contexts in which AI-generated developer actions and outputs can be tested consistently.

The role includes investigating integrations and documenting findings, as well as participating in calibration activities to keep technical grading and rubric application consistent.

  • Evaluate version control, pull requests, code review, issue tracking, and CI/CD workflows
  • Assess the technical correctness of AI-generated outputs
  • Verify results directly in development environments
  • Design and maintain repositories with realistic commit histories
  • Create CI workflows and reliable task contexts
  • Configure, document, and troubleshoot integrations and authentication flows
  • Systematically test undocumented product behavior and document findings
  • Participate in calibration activities and apply evaluation rubrics consistently

Required Qualifications

This role requires at least three years of professional software engineering experience and a strong command of Git and GitHub workflows. You should be comfortable diagnosing technical issues, authoring automation, and making careful judgments about developer-facing AI behavior.

Strong written and verbal English communication, attention to detail, process discipline, and sound technical judgment are essential.

  • At least three years of professional software engineering experience
  • Expertise with Git and GitHub, including branching, pull requests, code review, diffs, and CI log diagnostics
  • Proficiency with CI/CD systems, ideally GitHub Actions
  • Experience authoring custom workflows and supporting reproducible builds
  • Advanced Python or Bash scripting for environment setup and reset automation
  • Solid knowledge of REST APIs, OAuth, and webhook integrations
  • Ability to identify nuanced edge cases and apply rubrics consistently
  • Strong written and verbal English communication

Helpful Background

The following experience can further support success in the role. These areas are useful for managing realistic workspaces, evaluating AI agents, and recognizing subtle differences in technical behavior.

  • Experience with connectors across Claude, ChatGPT, AI assistants, or developer agent tools
  • Experience with Slack, Google Workspace, Microsoft 365, or cloud-workspace administration
  • Background in technical quality assurance, model evaluation, or data annotation
  • Familiarity with agent tool-calling or the Model Context Protocol

Work Schedule And Compensation

This worldwide contractor opportunity supports part-time work of 20 or more hours per week. Compensation is listed at USD $50-$70 per hour.

  • Remote and worldwide
  • 20+ hours per week
  • Part-time contractor engagement
  • USD $50-$70 hourly compensation

Build Your AI Training Career

AI training and data labeling are increasingly important ways to work in technology without leaving your professional specialty behind. By evaluating software engineering workflows, you can contribute to how advanced AI systems learn to develop, test, and troubleshoot software.

OpenTrain gives you a place to present credible experience, discover matching AI training opportunities, and build a longer-term portfolio around specialized technical work.

  • Apply through OpenTrain
  • Showcase your software engineering and evaluation experience
  • Build a portfolio in the fast-growing AI training industry

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Software Engineering AI Evaluation Expert

Use your software engineering judgment to evaluate AI-generated technical content, refine prompts, fact-check claims, and create high-quality engineering artifacts. Work remotely as a flexible contractor for 20+ hours per week.

Coding & Software
Document
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $100–$200/hr

Posted Aug 4, 2026

Software Engineering AI Code Evaluator

Use your software engineering expertise to create datasets, test AI-generated code, and assess solutions across the development lifecycle. This flexible contractor role offers 10 to 40 hours per week for candidates in six eligible countries.

Coding & Software
Computer Code Programming
Remote · United States, Canada, Austria +3 more
English
Part-time · Flexible
Entry level

Posted Aug 6, 2026

Software Engineering Code Evaluation Expert

Use your software engineering expertise to create coding challenges, reference solutions, and evaluations that improve AI systems. This remote contractor role offers flexible work of about 15 hours per week and listed pay of $50 to $100 per hour.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $50–$100/hr

Posted Aug 13, 2026