Join OpenTrain as a GPU Kernel Optimization Expert working 20+ hrs/week to profile, optimize, and document GPU kernels for AI projects using C++17, Python and CUDA/HIP; contractor pay USD $80–$100/hr, remote and worldwide.
Coding & Software
100% Remote Hourly · $80–$100/hr
$80–$100/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jul 13, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the centralized platform where people build careers in AI training and data labeling. We help freelancers consolidate proof-of-work, discover specialized projects, and grow a durable freelance portfolio focused on human-in-the-loop AI work.
As the hiring organization for this role, OpenTrain connects skilled practitioners to short- and medium-term AI training projects and gives contributors a single place to manage applications, deliverables, and career history.
About AI training and kernel work
AI training (data labeling, annotation, and human-feedback work) is the human side of building modern AI systems. Contributors shape how models behave by preparing, reviewing, and improving the examples and code those models learn from.
Kernel optimization is a specialized slice of that work: improving GPU kernel performance directly increases model throughput and efficiency, and your decisions will influence how production AI workloads run across modern GPU architectures.
The role — GPU Kernel Optimization Expert
You will evaluate and optimize GPU kernels used in AI projects: profile performance, find bottlenecks, change or suggest code-level improvements, and document why a change helps. This is contract, part-time work requiring 20+ hours per week.
Open to contributors worldwide who are fluent in English. Compensation is hourly in USD at $80–$100/hr. This role expects intermediate-level experience: at least one year of professional or graduate research experience working with GPUs.
Employment type: Contractor, Part-time
Time requirement: 20+ hours/week
Pay: USD $80–$100 per hour
Language: English (required); worldwide applicants welcome
What you'll do
Your day-to-day work centers on performance analysis, targeted code changes, and clear documentation of optimization decisions so other engineers can follow your reasoning.
Profile GPU kernels and interpret signals such as L2 cache hit rate, L2 throughput, and occupancy to guide improvements
Identify bottlenecks in kernel implementations and recommend or implement changes to improve utilization
Read and modify performance-sensitive code in C++17 and Python, and reason about lower-level GPU code
Apply CUDA, HIP, Slang/HLSL/GLSL, or other shader/kernel expertise to improve throughput and efficiency
Document optimization decisions and explain when and why specific profiler metrics are useful
Requirements
Candidates must be able to work independently on kernel-level performance tasks and produce clear, reproducible rationale for changes.
Strong command of core C++ up through C++17 for reading and modifying performance-critical code
Working knowledge of Python and Git for code review and lightweight tooling
Fluency in at least one GPU programming model (CUDA, HIP, Slang, HLSL, GLSL, etc.)
At least 1 year of professional or graduate-level research experience working with GPUs
Solid understanding of GPU profiler metrics and kernel optimization methods (occupancy, cache behavior, throughput)
Ability to reason about and optimize kernels without extensive prior context on every algorithm
Nice-to-have: experience with inline PTX, tensor-core optimization, CUDA C++ core libraries, NSight Compute, or open-source contributions in kernel optimization
Who should apply
This role fits engineers and researchers who enjoy low-level performance work and communicating technical trade-offs clearly. If you have practical GPU programming experience and a track record of profiling or speeding up kernels, you should apply.
You do not need domain expertise in every AI algorithm you touch — the role values strong performance reasoning and the ability to find generic bottlenecks across kernels.
Expert contractor role to design and optimize GPU software for LLM training (CUDA/WebGPU/GLSL). Remote, 20+ hrs/week, $60–$85/hr; requires advanced GPU programming and strong C++ for host-side integrations.
Join OpenTrain as a remote Performance Engineer focused on systems-level optimization, compiler engineering, and runtime performance for AI training data. This contract role pays $70–$110/hr and requires reliable full-time work (40 hours/week) from US-based contributors fluent in English.
Experienced ML engineers wanted for part-time, remote contract work evaluating production-grade model training, evaluation, and inference pipelines; requires 3+ years of ML engineering experience, strong Python, and PyTorch/TensorFlow/JAX familiarity. 20+ hours/week; apply through OpenTrain.