Evaluate how language models solve real software engineering problems using C#, GitHub repositories, Docker, and unit-test analysis. This remote contractor role requires 20 or more hours weekly and four hours of daily PST overlap.
Coding & Software
Remote
9 countries
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open to applicants in
India Pakistan Nigeria Kenya Egypt Ghana Bangladesh Türkiye Mexico
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. We help contributors discover specialized projects, build a professional AI training profile, and apply in minutes.
Creating an OpenTrain account is free. In this role, you will contribute directly to the human review and engineering work that helps improve modern AI systems.
About AI Training and Code Evaluation
AI training is the human side of building artificial intelligence. For coding systems, experienced engineers help create realistic programming tasks, assess model-generated solutions, and judge whether code works as expected.
This work combines practical software engineering with evaluation of large language models. Your reviews of repositories, issues, tests, and bug-fixing attempts can help shape how AI systems perform on real-world development problems.
The Role
OpenTrain AI is recruiting a C# Software Engineer for LLM code evaluation. You will help build and validate training datasets based on public GitHub repositories and repository histories, using human review to create verifiable software engineering tasks.
The role combines hands-on C# development, repository analysis, test-quality judgment, and practical assessment of how language models interact with real code. You may also lead or coordinate junior engineers when project needs require additional support.
Individual contractor assignment
Fully remote arrangement
Part-time options of 20, 30, or 40 hours per week
At least 20 hours per week required
Four hours of daily overlap with PST required
Available to candidates in India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico
What You'll Do
You will work with real-world open-source codebases and collaborate with researchers on the design of software engineering tasks and datasets. The work requires both implementation ability and careful judgment about software behavior and model performance.
Analyze and triage issues across trending open-source libraries.
Configure repositories and local development environments, including Dockerization and basic pipeline setup.
Evaluate unit-test coverage and overall test quality.
Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.
Identify repositories and issues that create meaningful challenges for language models.
Collaborate with researchers on task and dataset design.
Lead or coordinate junior engineers when project needs call for additional support.
Required Qualifications
This opportunity is listed as entry level in the project data, while the role itself requires at least three years of professional software engineering experience. Strong practical ability with C# and real-world codebases is essential.
At least three years of professional software engineering experience.
Strong experience with C#.
Proficiency with Git and Docker.
Practical experience with development environment configuration and basic software pipeline setup.
Ability to navigate complex codebases and work comfortably with real-world projects locally.
Ability to analyze and triage GitHub issues in open-source repositories.
Sound judgment when evaluating repository quality, unit-test coverage, test quality, expected software behavior, and LLM bug-fixing performance.
Helpful Background
Experience contributing to or evaluating open-source projects is useful. Previous participation in LLM research or evaluation work can also help you assess software engineering tasks from both implementation and model-performance perspectives.
Experience building or testing developer tools or automation agents may be valuable in this assignment.
Open-source project contribution or evaluation experience
Previous LLM research or evaluation work
Experience building or testing developer tools
Experience with automation agents
How This Work Fits Your Career
AI training and data-labeling work is a growing part of the technology industry. Contributors use their existing technical expertise to prepare examples, review outputs, and evaluate systems that are shaping the next generation of AI.
The remote, flexible structure can support a part-time workload alongside other professional or personal commitments while giving you hands-on experience with cutting-edge language model evaluation.
Remote work with a computer and internet connection
Flexible part-time scheduling within the required weekly commitment
Direct application of software engineering expertise to AI development
Opportunity to evaluate language models against realistic coding challenges
Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.
Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.
Use your C++ and software engineering expertise to evaluate how language models understand and fix real code. Work remotely with GitHub repositories, Docker, testing, and LLM evaluation through OpenTrain.