Use advanced CUDA, C++, GLSL, and WebGPU skills to optimize GPU kernels and evaluate performance across modern architectures. This remote contractor role offers $60-$100 per hour and 20+ hours per week.
Coding & Software
100% Remote Hourly · $60–$100/hr
$60–$100/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Aug 11, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized technical work that helps develop and evaluate modern AI systems.
Create a free OpenTrain account to build a profile, show your technical experience, discover relevant projects, and apply in minutes.
About AI Training and Technical Evaluation
AI training is the human side of building artificial intelligence. Alongside data annotation and human feedback, the field includes specialized programming and evaluation work that helps improve the systems, tools, and infrastructure used to develop AI.
This role focuses on GPU performance engineering and technical evaluation. Your analysis of kernels, shaders, and performance metrics can help assess practical approaches for high-performance computing workflows.
The Role
OpenTrain AI is seeking a CUDA GPU Kernel Optimization Engineer for remote contract work centered on GPU performance and real-world technical expertise. You will optimize GPU kernels, improve C++ and CUDA code, develop shader workflows, and evaluate GPU-based approaches through measurable performance analysis.
The role is listed as entry level, although the work requires demonstrated CUDA performance-tuning expertise, advanced C++, shader development experience, and proficiency with GPU profiling tools. Prior AI experience is not required.
Contractor and part-time engagement
Worldwide remote opportunity
20+ hours per week
Compensation of $60-$100 per hour
English-language work
What You'll Do
You will investigate GPU behavior, identify opportunities for improvement, and communicate optimization decisions clearly. The work combines hands-on programming with profiling, performance analysis, documentation, and technical collaboration.
Analyze, profile, and optimize GPU kernels to increase computational throughput on modern hardware.
Identify kernel bottlenecks and recommend targeted optimization strategies.
Refactor C++ and CUDA code for maintainability, efficiency, and adaptability across GPU architectures.
Implement GLSL and WebGPU shader logic and graphics or compute workflows.
Document optimization steps, findings, and performance improvements in clear technical reports.
Contribute to design discussions and evaluate GPU performance metrics and approaches.
Track developments in GPU programming and share relevant technical insights.
Required Skills and Experience
The strongest candidates will bring practical experience tuning GPU kernel performance and analyzing behavior across hardware generations. Clear technical communication is important because optimization findings and recommendations must be documented and discussed with others.
Demonstrated CUDA programming expertise and experience tuning GPU kernel performance.
Advanced C++ development skills in high-performance computing environments.
Hands-on GLSL and WebGPU experience for graphics or compute shader development.
Proficiency with GPU profiling tools such as Nsight, Visual Profiler, or comparable tools.
Strong analytical ability to reason about kernel performance across hardware generations.
Clear written and verbal technical communication.
Ability to produce clear technical documentation, reporting, and performance analysis.
Who Should Apply
Apply if your core strength is GPU programming and performance engineering. Prior AI experience is not required, so this opportunity may suit an engineer whose background is in high-performance computing, graphics, shader development, or related GPU optimization work.
Experience collaborating in remote, cross-disciplinary project settings is helpful. The work is designed for contributors who can manage a flexible part-time contractor schedule while delivering careful technical analysis.
GPU programmers with demonstrated kernel optimization experience
C++ engineers from high-performance computing environments
Engineers experienced with GLSL, WebGPU, graphics, or compute shaders
Technical contributors who communicate performance findings clearly
Candidates available for at least 20 hours per week
How to Apply Through OpenTrain
OpenTrain makes it easier to start and grow a career in AI training and data-labeling work, including specialized technical projects. Create a free account, build your profile around your GPU programming experience, and apply through OpenTrain.
This is a worldwide remote contractor opportunity with part-time availability and a listed rate of $60-$100 per hour. Review the role requirements carefully and highlight your CUDA optimization, C++, shader, profiling, and technical documentation experience.
Evaluate GPU and accelerator kernel tasks for correctness, performance, compilation, and runtime validity. This fully remote US contract role pays $70 to $90 per hour and requires expertise across CUDA, Triton, NKI, or Pallas.
Build and optimize GPU software for LLM training tasks using CUDA, WebGPU, GLSL, and C++. This expert-level contractor role offers worldwide remote work at $60 to $85 per hour for 20+ hours weekly.
Help evaluate NKI kernel development for advanced AI training, focusing on CUDA migration, Trainium performance, and numerical correctness. This US contract offers 20+ hours per week at $70-$90 per hour.