Data Engineering AI Evaluation Engineer
Build and validate Python data pipelines and benchmark tasks for advanced AI systems in a flexible, 3-month contractor role with 20+ hours per week.
Posted Jul 16, 2026
Review AI evaluation tasks, rubrics, verifiers, and outputs as a remote QC Engineer. This three-month contractor role requires 20+ hours weekly and is open to freelancers in nine countries.
Coding & Software
9 countries
Eligibility
Entry
Experience
Sep 21, 2026
Posted
Open to applicants in
OpenTrain AI is the hiring and contracting organization for this opportunity. OpenTrain helps people build careers in AI training and data labeling by connecting their skills, project experience, and professional profile to cutting-edge work.
AI data evaluation is the quality-focused work behind reliable artificial intelligence. Human experts review tasks, model outputs, scoring criteria, and verification methods to determine whether AI systems and the data used to assess them meet the intended standard.
OpenTrain is seeking an AI Data Evaluation QC Engineer to assess the quality and technical correctness of AI data evaluation projects. You will review tasks, outputs, evaluation criteria, verifiers, rubrics, and test cases while making independent quality and delivery judgments.
This opportunity is designed for a practicing software engineer who can work independently across quality control and technical delivery responsibilities. Experience with AI data, model evaluation, task generation, or similar projects is valuable.
You will help establish whether project work meets a demanding quality standard. The role requires independent validation rather than relying only on automated quality-control results, along with consistent ownership in demanding project environments.
You should be a practicing software engineer with strong technical fundamentals and sound technical judgment. You should understand quality-control pipelines, know what makes a task or verifier effective, and be able to assess technical correctness independently.
The role also calls for experience as a QC expert, technical project lead, delivery lead, or equivalent technical owner. Relevant experience may include designing, implementing, or reviewing quality frameworks at scale and owning final decisions for a technical project.
This is a fully remote, three-month contractor assignment. You must be available for at least 20 hours per week, including at least four hours per day and four hours of overlap with Pacific Time.
AI training and data labeling are expanding fields where people help shape how modern AI systems perform. OpenTrain gives contributors a place to build a credible record of this work, discover relevant opportunities, and develop a longer-term professional portfolio.
Keep exploring
Build and validate Python data pipelines and benchmark tasks for advanced AI systems in a flexible, 3-month contractor role with 20+ hours per week.
Posted Jul 16, 2026
Use your software QA expertise to evaluate AI-generated technical outputs, design edge-case tests, and deliver precise feedback. This remote contractor role pays $90-$175 per hour and requires prior paid AI evaluation experience.
Posted Sep 9, 2026
Create and validate challenging engineering simulation benchmarks that train and evaluate AI agents. This remote contractor role combines advanced engineering design, Python, open-source simulation, and model failure analysis.
Posted Sep 11, 2026
Browse related job pages
Expertise
Locations
Languages