Evaluate GPU and accelerator kernel tasks for correctness, performance, compilation, and runtime validity. This fully remote US contract role pays $70 to $90 per hour and requires expertise across CUDA, Triton, NKI, or Pallas.
Coding & Software
Remote Hourly · $70–$90/hr
$70–$90/hr
Compensation
1 country
Eligibility
Entry
Experience
Aug 27, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.
Creating an OpenTrain account is free, and this opportunity offers a way to apply specialized GPU programming expertise to the development and evaluation of advanced AI systems.
About AI Training Work
AI training is the human side of building artificial intelligence. Specialists review code, assess model outputs, and provide structured feedback that helps AI systems become more capable, reliable, and useful.
In this role, your technical evaluations will support training and evaluation workflows involving GPU and accelerator kernel development. The work is fully remote and combines deep engineering expertise with cutting-edge AI development.
The GPU Kernel Evaluation Expert Role
OpenTrain AI is seeking a GPU Kernel Evaluation Expert to assess the quality, correctness, and completeness of GPU and accelerator kernel development tasks used to train and evaluate frontier AI models.
You will evaluate numerical correctness, benchmarking fairness, task scope, compilation validity, and runtime behavior across a range of kernel task types. Each submission requires clear, rubric-based written feedback.
Fully remote contract role for eligible candidates in the United States
Pay range: $70.00 to $90.00 per hour
The role description specifies a 40-hour-per-week commitment
The structured time requirement is listed as 20+ hours per week
Employment types: contractor and part-time
Primary language: English
What You'll Evaluate
You will review technical submissions and task designs across multiple aspects of GPU and accelerator kernel development. Your assessments should be accurate, consistent, and grounded in the applicable evaluation rubric.
GPU and accelerator kernel tasks for quality, correctness, and completeness
Numerical correctness using absolute, relative, and ULP tolerances
Selection and suitability of reference implementations
Performance-benchmarking fairness and profiling results
Compilation and runtime validity across different environments
Generation from specification, translation or lowering, migration, debugging, optimization, and operator-fusion tasks
Written, rubric-based feedback for every evaluated task
Required Qualifications
The listing is marked entry level, but the role specifically requires at least three years of hands-on experience developing, optimizing, or verifying GPU or accelerator kernels. You must have experience in at least two of CUDA, Triton, NKI, or Pallas for JAX.
3+ years developing, optimizing, or verifying GPU or accelerator kernels
Hands-on experience with at least two of CUDA, Triton, NKI, or Pallas for JAX
Strong understanding of numerical-correctness criteria for kernels
Experience with Nsight, NCU, roofline analysis, or framework-native profiling tools
Familiarity with common compilation and runtime failure modes
Experience with at least three task types: generation, translation or lowering, migration, debugging, optimization, or operator fusion
Helpful Technical Background
The following experience is helpful for evaluating a broad range of kernel tasks and accelerator environments. These qualifications are preferred background rather than listed minimum requirements.
Experience across NVIDIA GPU ecosystems such as CUDA and Triton
Experience with custom-accelerator ecosystems such as NKI, Pallas, or TPU
Compiler engineering, MLIR, or intermediate-representation lowering
Memory-hierarchy optimization, including shared-memory tiling, register pressure, bank conflicts, and coalescing patterns
Contributions to cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
How This Work Supports AI Development
Modern AI systems depend on people who can inspect technical outputs, identify failure modes, and distinguish reliable results from misleading ones. By evaluating kernel implementations and their benchmarks, you help improve the quality of the data and feedback used in AI development.
This is a specialized path within the broader AI-training industry, where technical contributors can use software and systems expertise to shape how advanced models are built and evaluated.
Apply Through OpenTrain
If your background matches the kernel development and evaluation requirements, create a free OpenTrain account and apply through the platform. Review the stated schedule, location eligibility, and technical expectations before submitting your application.
Confirm eligibility to work remotely from the United States
Highlight your CUDA, Triton, NKI, or Pallas experience
Describe your profiling, benchmarking, and numerical-validation work
Show experience across at least three listed kernel task types
Use advanced CUDA, C++, GLSL, and WebGPU skills to optimize GPU kernels and evaluate performance across modern architectures. This remote contractor role offers $60-$100 per hour and 20+ hours per week.
Build and optimize GPU software for LLM training tasks using CUDA, WebGPU, GLSL, and C++. This expert-level contractor role offers worldwide remote work at $60 to $85 per hour for 20+ hours weekly.
Join OpenTrain as a Machine Learning Engineer evaluating real-world ML systems through benchmark-driven coding tasks. Work remotely on training, inference, debugging, and model evaluation for at least 20 hours per week.