Skip to content
OpenTrain AIFor AI Companies

Trainium NKI Kernel Expert

Help evaluate NKI kernel development for advanced AI training, focusing on CUDA migration, Trainium performance, and numerical correctness. This US contract offers 20+ hours per week at $70-$90 per hour.

OpenTrain AI

Coding & Software

Remote Hourly · $70–$90/hr

$70–$90/hr

Compensation

1 country

Eligibility

Entry

Experience

Aug 28, 2026

Posted

Open to applicants in

United States

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is hiring contractors for specialized AI training work. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build a professional profile, and apply in minutes.

Creating an OpenTrain account is free, and this role offers an opportunity to contribute directly to the evaluation processes that help improve advanced AI systems.

About AI Training and Code Evaluation

AI training is the human side of building artificial intelligence. People with technical expertise review code, assess model outputs, and provide structured feedback that helps AI systems become more accurate, capable, and reliable.

In this role, your evaluation of NKI kernel development tasks will help improve the quality of training data used for advanced AI model development.

The Role

OpenTrain is recruiting a Trainium NKI Kernel Expert to evaluate the quality and correctness of NKI development tasks used to train advanced AI models. The work centers on CUDA-to-NKI migration fidelity, Trainium-specific performance optimization, and cross-platform numerical accuracy.

You will deliver clear, rubric-based written feedback on each task. The role is listed as entry level, while the technical requirements call for at least two years of hands-on NKI kernel development experience.

  • Contractor and part-time opportunity
  • US-based role
  • 20+ hours per week
  • English-language work
  • $70-$90 per hour

What You'll Do

You will assess technical implementation quality across NKI kernel development tasks and determine whether each solution is correct, appropriate for the target hardware, and aligned with defined evaluation standards.

  • Evaluate NKI kernel development tasks for quality, correctness, and hardware appropriateness on AWS Trainium and Inferentia2.
  • Assess CUDA-to-NKI migration fidelity and identify performance bottlenecks.
  • Define and apply cross-platform numerical-correctness standards for GPU and Trainium implementations.
  • Provide clear, detailed, rubric-based written feedback on each task.

Required Qualifications

This role requires specialized experience developing and evaluating kernels for AWS Trainium and Inferentia2. Candidates should be comfortable analyzing low-level computation, memory movement, profiling data, migration quality, and numerical behavior across hardware platforms.

  • At least 2 years of hands-on NKI kernel development targeting AWS Trainium or Inferentia2.
  • Strong understanding of NKI tile-based computation, SBUF, PSUM, HBM memory hierarchy, partition-dimension constraints, and DMA orchestration.
  • Demonstrated experience assessing CUDA-to-NKI migration quality.
  • Familiarity with Trainium-specific profiling, including NeuronCore pipeline utilization, tensor-engine throughput, and memory bandwidth.
  • Experience defining or evaluating GPU versus Trainium numerical-correctness standards.

Helpful Technical Background

The following experience is helpful for evaluating a broad range of NKI and Trainium tasks, but is listed as additional background rather than a required qualification.

  • Experience with the AWS Neuron SDK or Neuron Compiler internals.
  • Contributions to NKI kernel libraries.
  • Prior CUDA or Triton kernel development.
  • Familiarity with NeuronCore-v2 architecture and supported FP32, BF16, FP8, and INT8 data types.
  • Experience benchmarking machine learning training workloads on Trn1 or Trn2 instances.

Why This Work Matters

AI training and data-labeling work is a rapidly growing part of the technology industry. Technical reviewers help shape how modern AI systems perform by checking examples, code, and model behavior against rigorous standards.

This opportunity is especially suited to an expert who wants flexible, part-time work while contributing to cutting-edge AI development through detailed technical evaluation.

How to Apply

Create a free OpenTrain account, build your technical profile, and apply in minutes. Be prepared to highlight your NKI kernel development experience, CUDA-to-NKI migration work, Trainium profiling knowledge, and experience evaluating numerical correctness.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

GPU Kernel Evaluation Expert

Evaluate GPU and accelerator kernel tasks for correctness, performance, compilation, and runtime validity. This fully remote US contract role pays $70 to $90 per hour and requires expertise across CUDA, Triton, NKI, or Pallas.

Coding & Software
Computer Code Programming
Remote · United States
English
Part-time · Flexible
Entry level
Hourly · $70–$90/hr

Posted Aug 27, 2026

Operating Systems Expert AI Training

Use your Operating Systems expertise to create, review, and refine technical AI training data. This India-based contractor role offers flexible work under 20 hours per week at $25 per hour.

Coding & Software
Computer Code Programming
Remote · India
English
Part-time · Flexible
Entry level
Hourly · $25/hr

Posted Jan 2, 2025

CUDA GPU Kernel Optimization Engineer

Use advanced CUDA, C++, GLSL, and WebGPU skills to optimize GPU kernels and evaluate performance across modern architectures. This remote contractor role offers $60-$100 per hour and 20+ hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Hourly · $60–$100/hr

Posted Aug 11, 2026