Build realistic coding-agent benchmarks, test suites, and security-focused evaluators for large language models. This part-time contractor role is open worldwide and requires senior production engineering experience.
Coding & Software
100% Remote
Worldwide
Eligibility
Entry
Experience
Aug 6, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized projects where technical experts help evaluate and improve cutting-edge artificial intelligence.
Contractor opportunity with OpenTrain AI
Open worldwide
Part-time commitment of 20+ hours per week
English-language work
About AI Training and LLM Evaluation
AI training is the human side of building modern artificial intelligence. Experts create examples, review model outputs, and test whether AI systems follow instructions safely and accurately. In this role, your software engineering and security expertise will help measure whether coding agents can complete useful tasks without violating system constraints or developer intent.
Work at the intersection of software engineering and AI evaluation
Help assess coding agents using realistic repositories and runtime scenarios
Contribute to how advanced AI systems behave in production-like environments
The Role
OpenTrain is recruiting a Senior Software Engineer focused on LLM evaluation and repository validation. You will design and build realistic coding-agent benchmark tasks using production-like repositories, tests, configurations, documentation, and runtime scenarios.
The work combines software engineering with AI evaluation. Benchmark outcomes must measure task completion while preserving data integrity, privacy, permissions, oversight mechanisms, and other system constraints. The listing identifies this as an entry-level opportunity, while the required background includes at least eight years of hands-on software engineering experience.
Subject matter: Coding Agent Benchmark Evaluation
Data type: Computer code and programming
Work types: Programming, evaluation and rating, and red teaming
No specific hourly or fixed project rate is provided
What You'll Do
You will create benchmark tasks that reflect realistic engineering work and develop the materials needed to run, score, and calibrate them. You will also analyze coding-agent behavior and work with engineering, quality, and security stakeholders to improve evaluator reliability and benchmark difficulty.
Create benign tasks for bug fixes, feature additions, CI repairs, configuration migrations, integration updates, and runtime-state fixes
Define utility and safety requirements that make benchmark outcomes clear and testable
Develop visible and hidden test suites for task completion and unsafe behavior
Produce safe and intentionally unsafe reference solutions that expose alignment failures
Evaluate shortcuts such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation
Package runnable repositories with prompts, metadata, evaluators, reference patches, scoring rubrics, and calibration notes
Analyze coding-agent rollouts and improve evaluator reliability and benchmark difficulty
Required Qualifications
This role requires a strong production engineering background. Candidates must be able to build and maintain software used by real customers, diagnose distributed system failures, and identify and remediate security vulnerabilities.
Bachelor's or master's degree in computer science or a related technical discipline
At least eight years of hands-on software engineering experience in leading product companies or technology startups
Strong Python programming experience plus experience with Java, Go, or C++
Experience building and maintaining production software used by real customers
Strong knowledge of software architecture, debugging, performance optimization, deployment pipelines, monitoring, logging, and observability
Ability to diagnose failures using logs, metrics, and distributed tracing
Understanding of high availability, fault tolerance, scalability, and disaster recovery
Experience identifying vulnerabilities through code review or security assessment
Experience implementing secure fixes and validating remediation
Understanding of coding-agent benchmark design, test evaluation, and software reliability
Helpful Security Background
Security-focused evaluation benefits from familiarity with common software vulnerabilities and failure modes. This knowledge supports the creation of realistic unsafe scenarios and reliable tests for secure remediation.
Injection attacks
Authentication and authorization flaws
Memory safety
Race conditions
Deserialization
Secrets management
Input validation
Why Build Your AI Training Career With OpenTrain
AI training and data-labeling work is a fast-growing way to contribute to technology. OpenTrain helps freelancers manage specialized opportunities, build a credible AI training portfolio, and find projects that match their skills.
Work remotely from anywhere with an internet connection
Choose flexible part-time work that fits your schedule
Showcase technical experience through an AI training portfolio
Discover opportunities that can support a longer-term career in AI training
How to Apply
Create a free OpenTrain account and apply through the platform. Be prepared to demonstrate your senior software engineering background, production systems experience, Python expertise, distributed debugging skills, and security evaluation experience.
Worldwide contractor opportunity
Part-time schedule of 20+ hours per week
English required for project communication and evaluation materials
Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.
Evaluate how AI models fix real software bugs in open-source Python repositories. This flexible, part-time contractor role focuses on GitHub issue triage, Docker environments, testing, and LLM performance assessment.
Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.