Skip to content
OpenTrain AIFor AI Companies

Senior Software Engineer, LLM Evaluation

Build realistic coding-agent benchmarks, test suites, and security-focused evaluators for large language models. This part-time contractor role is open worldwide and requires senior production engineering experience.

OpenTrain AI

Coding & Software

100% Remote

Worldwide

Eligibility

Entry

Experience

Aug 6, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized projects where technical experts help evaluate and improve cutting-edge artificial intelligence.

  • Contractor opportunity with OpenTrain AI
  • Open worldwide
  • Part-time commitment of 20+ hours per week
  • English-language work

About AI Training and LLM Evaluation

AI training is the human side of building modern artificial intelligence. Experts create examples, review model outputs, and test whether AI systems follow instructions safely and accurately. In this role, your software engineering and security expertise will help measure whether coding agents can complete useful tasks without violating system constraints or developer intent.

  • Work at the intersection of software engineering and AI evaluation
  • Help assess coding agents using realistic repositories and runtime scenarios
  • Contribute to how advanced AI systems behave in production-like environments

The Role

OpenTrain is recruiting a Senior Software Engineer focused on LLM evaluation and repository validation. You will design and build realistic coding-agent benchmark tasks using production-like repositories, tests, configurations, documentation, and runtime scenarios.

The work combines software engineering with AI evaluation. Benchmark outcomes must measure task completion while preserving data integrity, privacy, permissions, oversight mechanisms, and other system constraints. The listing identifies this as an entry-level opportunity, while the required background includes at least eight years of hands-on software engineering experience.

  • Subject matter: Coding Agent Benchmark Evaluation
  • Data type: Computer code and programming
  • Work types: Programming, evaluation and rating, and red teaming
  • No specific hourly or fixed project rate is provided

What You'll Do

You will create benchmark tasks that reflect realistic engineering work and develop the materials needed to run, score, and calibrate them. You will also analyze coding-agent behavior and work with engineering, quality, and security stakeholders to improve evaluator reliability and benchmark difficulty.

  • Create benign tasks for bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes
  • Define utility and safety requirements that make benchmark outcomes clear and testable
  • Develop visible and hidden test suites for task completion and unsafe behavior
  • Produce safe and intentionally unsafe reference solutions that expose alignment failures
  • Evaluate shortcuts such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation
  • Package runnable repositories with prompts, metadata, evaluators, reference patches, scoring rubrics, and calibration notes
  • Analyze coding-agent rollouts and improve evaluator reliability and benchmark difficulty

Required Qualifications

This role requires a strong production engineering background. Candidates must be able to build and maintain software used by real customers, diagnose distributed system failures, and identify and remediate security vulnerabilities.

  • Bachelor's or master's degree in computer science or a related technical discipline
  • At least eight years of hands-on software engineering experience in leading product companies or technology startups
  • Strong Python programming experience plus experience with Java, Go, or C++
  • Experience building and maintaining production software used by real customers
  • Strong knowledge of software architecture, debugging, performance optimization, deployment pipelines, monitoring, logging, and observability
  • Ability to diagnose failures using logs, metrics, and distributed tracing
  • Understanding of high availability, fault tolerance, scalability, and disaster recovery
  • Experience identifying vulnerabilities through code review or security assessment
  • Experience implementing secure fixes and validating remediation
  • Understanding of coding-agent benchmark design, test evaluation, and software reliability

Helpful Security Background

Security-focused evaluation benefits from familiarity with common software vulnerabilities and failure modes. This knowledge supports the creation of realistic unsafe scenarios and reliable tests for secure remediation.

  • Injection attacks
  • Authentication and authorization flaws
  • Memory safety
  • Race conditions
  • Deserialization
  • Secrets management
  • Input validation

Why Build Your AI Training Career With OpenTrain

AI training and data-labeling work is a fast-growing way to contribute to technology. OpenTrain helps freelancers manage specialized opportunities, build a credible AI training portfolio, and find projects that match their skills.

  • Work remotely from anywhere with an internet connection
  • Choose flexible part-time work that fits your schedule
  • Showcase technical experience through an AI training portfolio
  • Discover opportunities that can support a longer-term career in AI training

How to Apply

Create a free OpenTrain account and apply through the platform. Be prepared to demonstrate your senior software engineering background, production systems experience, Python expertise, distributed debugging skills, and security evaluation experience.

  • Worldwide contractor opportunity
  • Part-time schedule of 20+ hours per week
  • English required for project communication and evaluation materials
  • Apply through OpenTrain

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

Senior Software Engineer LLM Evaluation

Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

Senior Python Software Engineer LLM Evaluation

Evaluate how AI models fix real software bugs in open-source Python repositories. This flexible, part-time contractor role focuses on GitHub issue triage, Docker environments, testing, and LLM performance assessment.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Jul 16, 2026

Senior C++ Software Engineer for LLM Evaluation

Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +6 more
English
Part-time · Flexible
Entry level

Posted Aug 30, 2026