Design and evaluate realistic coding tasks that train and test large language models for software engineering. Remote, contractor role (20+ hrs/week) for experienced engineers who can write tasks, tests, and review AI-generated code.
Coding & Software
100% Remote
Worldwide
Eligibility
Expert
Experience
Jul 16, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform where people build careers in AI training and data labeling. Creating an OpenTrain account is free; contributors discover projects, build a unified portfolio, and grow into durable freelance careers teaching AI.
Why AI training matters
AI training (data labeling and human feedback) is the human side of building modern AI. Engineers and annotators create examples, evaluate outputs, and shape how models behave — making this work a direct way to influence state-of-the-art systems while working remotely and flexibly.
The role
You will design software engineering tasks and evaluate model-generated code to improve LLM performance on realistic engineering problems. This contractor, part-time role expects 20+ hours per week and is open worldwide; work is conducted in English.
Position type: Contractor, part-time (20+ hours/week).
Location: Remote — open to contributors worldwide who work in English.
What you'll do
Design realistic software engineering tasks that test reasoning and coding ability.
Write clear task descriptions, metadata, solution explanations, and validation test logic.
Evaluate AI-generated code for correctness, maintainability, and engineering best practices.
Create scenarios for feature implementation, bug fixing, refactoring, migrations, and code porting.
Work across repositories in multiple programming languages and frameworks.
Analyze model failures and provide structured feedback to improve future outputs.
Required qualifications
3+ years of professional software engineering experience.
Proficiency in one or more of: Python, Java, C/C++, JavaScript/TypeScript, Go, PHP, Ruby, or Rust.
Experience with modern development workflows: feature implementation, APIs, networking, migrations, or large codebases.
Strong understanding of data structures, algorithms, and system design.
Ability to write clear technical documentation and English explanations.
Proven experience reviewing code for correctness, maintainability, and engineering best practices.
Demonstrated ability to design realistic coding tasks and validate solutions with test logic.
Helpful background
Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field.
Hands-on experience with web applications, REST APIs, or networking software.
Strong analytical, debugging, and problem-solving skills.
Who should apply
Experienced software engineers who enjoy crafting realistic problems, writing tests, and mentoring models through high-quality feedback. Ideal candidates combine practical coding experience with strong written communication and a desire to shape how LLMs perform on real engineering tasks.
Join OpenTrain as a remote contractor to evaluate LLM performance on real open-source codebases using Ruby, Git, and Docker. This part-time role requires at least 20 hours/week, a 4-hour PST overlap, and candidates based in specified countries.
Join OpenTrain as a remote contractor building and evaluating LLM performance on real C++ codebases; flexible 20/30/40 hr/week schedules and opportunities to lead junior engineers. Work with researchers to design verifiable engineering tasks, triage issues, run code, and rate model outputs.
Join OpenTrain to build LLM evaluation and training datasets by validating real open-source codebases with C#. This remote contractor role requires 3+ years of software engineering, 20+ hours/week (options to 30–40) and is open to candidates in specified countries.