Skip to content
OpenTrain AIFor AI Companies

Databricks Specialist — Spark With Python/Java/Scala

Join OpenTrain AI as a remote Databricks Specialist working 20+ hours/week to design and optimize large-scale Spark data pipelines; pay is USD $12/hr and you'll be asked to document your language experience and weekly availability. Candidates must have at least 5 years of hands-on Databricks and Spa

OpenTrain AI

Coding & Software

100% Remote Hourly · $12/hr

$12/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Nov 12, 2024

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We connect skilled contributors with hands-on projects that shape how modern AI systems behave.

We hire contractors directly: when you join this role you’ll be contracting with OpenTrain AI, working remotely and contributing to real engineering work that improves how models are built and maintained.

About AI Training and Data Engineering Work

AI training includes many tasks that help models learn from human-created and human-reviewed data. This role sits at the data-engineering side of that work: building, testing and optimizing data processing systems that feed AI pipelines.

This kind of work is remote, flexible, and accessible to experienced engineers who enjoy designing reliable ETL, debugging large codebases, and improving performance at scale.

The Role

We are hiring experienced data engineers to work as Databricks Specialists. You will design and optimize large-scale data workflows, build and tune Spark jobs, and help keep Databricks pipelines performant and reliable.

This is a contractor, part-time opportunity (20+ hours/week) paid at USD $12 per hour. The role is fully remote and open worldwide; strong written English (B1 or B2) is required for collaboration and documentation.

  • Employment type: Contractor, Part-time
  • Time requirement: 20+ hours/week
  • Pay: USD $12 per hour
  • Location: Remote, worldwide (English B1/B2 required)

What You'll Do

You will build and maintain Databricks-based data pipelines and ETL, analyze and optimize Spark jobs, and debug issues in large distributed data processing systems.

Work includes hands-on coding, performance tuning, troubleshooting OutOfMemory and other stability problems, documenting fixes, and collaborating asynchronously with a remote engineering team.

  • Develop efficient Databricks workflows and ETL pipelines
  • Optimize Spark jobs for performance and memory usage
  • Analyze, debug, and test large codebases in Databricks
  • Write clear documentation and collaborate remotely

Requirements

Do not apply unless you meet the concrete, stated requirements below – we will ask for specifics during screening.

You must be able to list the exact number of years of experience you have with each programming language you know and state how many hours per week you are available for this project.

  • Minimum 5 years of hands-on experience working with Databricks
  • Deep expertise with Apache Spark, including building, optimizing, and troubleshooting Spark-based systems
  • 5+ years of experience in at least one of: Python, Java, SQL, Scala, or Spark; indicate which language(s) and years for each
  • Strong experience building and optimizing data pipelines and ETL processes
  • Experience analyzing, debugging, and testing large code bases and navigating complex documentation
  • Familiarity with cloud platforms such as Azure or AWS (preferred)
  • Ability to work independently and solve complex technical problems
  • Excellent communication skills for remote collaboration
  • English proficiency: B1 or B2

Interview, Tests, and What We’ll Ask

The interview includes technical screening and two written test questions below. Please be ready to answer them fully; do not conclude the live interview until both are answered. When you apply, include a brief interview summary that lists the programming language(s) you know and the exact years of experience you have with each, plus your weekly availability in hours.

We will evaluate your answers for correctness and completeness and use them to score the interview.

  • Before interview: include each language you know and years of experience for that language, and your available hours/week
  • Test Question 1 — Databricks Debugging and Optimization: You are given a PySpark job in Databricks that processes a large dataset but keeps failing with an OutOfMemoryError. What steps would you take to debug and resolve this issue? Provide specific adjustments or optimizations you would apply in Da
  • Test Question 2 — Code Review and Documentation: Review this Python code in a Databricks notebook: data = spark.read.csv("/path/to/file.csv", header=True) filtered_data = data.filter(data["column"] > 100) result = filtered_data.groupBy("category").count() result.show() Identify issues or areas for i

How to Apply and Next Steps

Apply through OpenTrain by submitting your resume/CV and a short cover note that lists (1) every programming language you know and the exact years of experience for each, and (2) how many hours per week you can commit. Include any Databricks or Spark references you can share.

If selected, you will be invited to a live technical interview where you must answer the two test questions above. We will score technical accuracy and completeness and confirm availability before offering a contract.

  • Include language-by-language years of experience in your application
  • State your weekly availability in hours and confirm acceptance of USD $12/hr
  • Be prepared to answer both test questions during the live interview

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

Data Analytics & Visualization Specialist (Python, Dash)

Part-time contractor role building ETL pipelines and interactive dashboards using Python, Plotly/Dash, and SQL. Entry-level, remote work under 20 hrs/week at $25/hr — ideal for hands-on data analysts who document clearly and translate questions into visuals.

Coding & Software
Computer Code Programming
Remote · Worldwide
Part-time · Flexible
Entry level
Hourly · $25/hr

Posted Sep 3, 2025

Data Scientist — Mathematical Statistics (Python)

Join OpenTrain AI as a part-time contractor applying mathematical statistics with Python (numpy, scipy, statsmodels, pandas) to analyze messy datasets, run hypothesis tests and regressions, and produce clear, reproducible summaries; remote, <20 hrs/week at $25/hr.

Coding & Software
Computer Code Programming
Remote · Worldwide
Part-time · Flexible
Entry level
Hourly · $25/hr

Posted Sep 3, 2025

Vibecode Specialist - Web Scraping & Data Extraction

Join OpenTrain AI to build end-to-end Python scraping pipelines and deliver clean structured datasets (CSV/JSON/Sheets). Remote, 20+ hrs/week at $20/hr; you must have hands-on experience with dynamic/JS-rendered sites and B2+ English.

Coding & Software
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $20/hr

Posted Feb 26, 2026