Skip to content
OpenTrain AIFor AI Companies

CUDA GPU Kernel Optimization Engineer

Optimize GPU kernels and improve C++ and CUDA code in a remote contractor role paying $60-$100 per hour. You will also build shader workflows and analyze performance across GPU hardware.

Apply now
OpenTrain AI

Coding & Software

Remote Hourly · $60–$100/hr

$60–$100/hr

Compensation

230 countries

Eligibility

Entry

Experience

Aug 11, 2026

Posted

Open to applicants in

United States
+ more
  • Åland Islands
  • Albania
  • Algeria
  • American Samoa
  • Andorra
  • Angola
  • Anguilla
  • Antarctica
  • Antigua & Barbuda
  • Argentina
  • Armenia
  • Aruba
  • Australia
  • Austria
  • Azerbaijan
  • Bahamas
  • Bahrain
  • Bangladesh
  • Barbados
  • Belgium
  • Belize
  • Benin
  • Bermuda
  • Bhutan
  • Bolivia
  • Bosnia & Herzegovina
  • Botswana
  • Bouvet Island
  • Brazil
  • British Indian Ocean Territory
  • British Virgin Islands
  • Brunei
  • Bulgaria
  • Burkina Faso
  • Burundi
  • Cambodia
  • Cameroon
  • Canada
  • Cape Verde
  • Caribbean Netherlands
  • Cayman Islands
  • Central African Republic
  • Chad
  • Chile
  • Christmas Island
  • Cocos (Keeling) Islands
  • Colombia
  • Comoros
  • Congo - Brazzaville
  • Cook Islands
  • Costa Rica
  • Côte d’Ivoire
  • Croatia
  • Curaçao
  • Cyprus
  • Czechia
  • Denmark
  • Djibouti
  • Dominica
  • Dominican Republic
  • Ecuador
  • Egypt
  • El Salvador
  • Equatorial Guinea
  • Eritrea
  • Estonia
  • Eswatini
  • Ethiopia
  • Falkland Islands (Islas Malvinas)
  • Faroe Islands
  • Fiji
  • Finland
  • France
  • French Guiana
  • French Polynesia
  • French Southern Territories
  • Gabon
  • Gambia
  • Georgia
  • Germany
  • Ghana
  • Gibraltar
  • Greece
  • Greenland
  • Grenada
  • Guadeloupe
  • Guam
  • Guatemala
  • Guernsey
  • Guinea
  • Guinea-Bissau
  • Guyana
  • Haiti
  • Heard & McDonald Islands
  • Honduras
  • Hungary
  • Iceland
  • India
  • Indonesia
  • Ireland
  • Isle of Man
  • Israel
  • Italy
  • Jamaica
  • Japan
  • Jersey
  • Jordan
  • Kazakhstan
  • Kenya
  • Kiribati
  • Kosovo
  • Kuwait
  • Kyrgyzstan
  • Laos
  • Latvia
  • Lebanon
  • Lesotho
  • Liberia
  • Liechtenstein
  • Lithuania
  • Luxembourg
  • Madagascar
  • Malawi
  • Malaysia
  • Maldives
  • Mali
  • Malta
  • Marshall Islands
  • Martinique
  • Mauritania
  • Mauritius
  • Mayotte
  • Mexico
  • Micronesia
  • Moldova
  • Monaco
  • Mongolia
  • Montenegro
  • Montserrat
  • Morocco
  • Mozambique
  • Namibia
  • Nauru
  • Nepal
  • Netherlands
  • New Caledonia
  • New Zealand
  • Nicaragua
  • Niger
  • Nigeria
  • Niue
  • Norfolk Island
  • North Macedonia
  • Northern Mariana Islands
  • Norway
  • Oman
  • Pakistan
  • Palau
  • Palestine
  • Panama
  • Papua New Guinea
  • Paraguay
  • Peru
  • Philippines
  • Pitcairn Islands
  • Poland
  • Portugal
  • Puerto Rico
  • Qatar
  • Réunion
  • Romania
  • Rwanda
  • Samoa
  • San Marino
  • São Tomé & Príncipe
  • Saudi Arabia
  • Senegal
  • Serbia
  • Seychelles
  • Sierra Leone
  • Singapore
  • Sint Maarten
  • Slovakia
  • Slovenia
  • Solomon Islands
  • South Africa
  • South Georgia & South Sandwich Islands
  • South Korea
  • Spain
  • Sri Lanka
  • St. Barthélemy
  • St. Helena
  • St. Kitts & Nevis
  • St. Lucia
  • St. Martin
  • St. Pierre & Miquelon
  • St. Vincent & Grenadines
  • Suriname
  • Svalbard & Jan Mayen
  • Sweden
  • Switzerland
  • Taiwan
  • Tajikistan
  • Tanzania
  • Thailand
  • Timor-Leste
  • Togo
  • Tokelau
  • Tonga
  • Trinidad & Tobago
  • Tunisia
  • Türkiye
  • Turkmenistan
  • Turks & Caicos Islands
  • Tuvalu
  • U.S. Outlying Islands
  • U.S. Virgin Islands
  • Uganda
  • United Arab Emirates
  • United Kingdom
  • United States
  • Uruguay
  • Uzbekistan
  • Vanuatu
  • Vatican City
  • Vietnam
  • Wallis & Futuna
  • Western Sahara
  • Zambia
  • Zimbabwe

The work

You will optimize GPU kernels to improve computational throughput on modern hardware. The work includes profiling performance, finding bottlenecks, and recommending focused changes.

You will also improve C++ and CUDA code, develop GLSL and WebGPU shader workflows, document your findings, and help evaluate GPU performance approaches through measurable results.

  • Analyze, profile, and optimize GPU kernels.
  • Identify kernel bottlenecks and recommend optimization strategies.
  • Refactor C++ and CUDA code for efficiency, maintainability, and use across GPU architectures.
  • Implement GLSL and WebGPU graphics or compute shader logic.
  • Write clear technical reports covering optimization steps and performance improvements.
  • Contribute to design discussions and evaluate GPU performance metrics.
  • Track developments in GPU programming and share useful technical insights.

What it pays and takes

This is a remote, part-time contractor role. The listing is marked entry level, but the work requires demonstrated GPU programming and performance-engineering expertise.

English is required. The role is available in the countries listed for this posting.

  • Pay: $60-$100 per hour.
  • Time: 20 or more hours per week.
  • Work type: Remote contractor and part-time.
  • Language: English.
  • Required: CUDA programming and GPU kernel performance tuning.
  • Required: Advanced C++ development in high-performance computing environments.
  • Required: Hands-on GLSL and WebGPU experience for graphics or compute shaders.
  • Required: Experience with Nsight, Visual Profiler, or comparable GPU profiling tools.
  • Required: Strong analysis of kernel performance across hardware generations.
  • Required: Clear written and verbal technical communication.
  • Helpful: Experience working remotely with cross-disciplinary teams.
  • Prior AI experience is not required.

How it works

Apply on OpenTrain with your resume, then complete the application on the hiring site.

About AI training work

AI training work uses human expertise to build and evaluate the code, examples, and feedback that help AI systems perform better. OpenTrain hires and contracts contributors for this work, and experienced technical specialists help assess and improve the systems behind modern AI tools.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

GPU Programming Software Engineer

Build and optimize GPU software for LLM training tasks as an expert contractor. Work 20 or more hours per week for $60 to $95 per hour using CUDA, WebGPU, GLSL, and C++.

Coding & Software
Computer Code Programming
Remote · Andorra, United Arab Emirates, Antigua & Barbuda +227 more
English
Part-time · Flexible
Expert level
Hourly · $60–$95/hr

Posted Jul 29, 2026

DevOps Engineer GPU and LLM Infrastructure

Build and operate GPU infrastructure and production LLM serving systems on Google Cloud. This remote contractor role requires Python, Kubernetes, Docker, GPU operations, and model-serving experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Sep 11, 2026

Machine Learning Engineer Benchmark Evaluation

Build and evaluate machine learning training, inference, and benchmarking pipelines as a remote contractor. This role requires 3+ years of ML engineering experience, strong Python skills, and availability for 20+ hours each week.

Coding & Software
Computer Code Programming
Remote · India, Pakistan, Nigeria +7 more
English
Part-time · Flexible
Intermediate level

Posted Jul 16, 2026