Databricks Specialist — Spark With Python/Java/Scala
Join OpenTrain AI as a remote Databricks Specialist working 20+ hours/week to design and optimize large-scale Spark data pipelines; pay is USD $12/hr and you'll be asked to document your language experience and weekly availability. Candidates must have at least 5 years of hands-on Databricks and Spa
Coding & Software
100% Remote Hourly · $12/hr
$12/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Nov 12, 2024
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the #1 platform for finding and building careers in AI training and data labeling. We connect skilled contributors with hands-on projects that shape how modern AI systems behave.
We hire contractors directly: when you join this role you’ll be contracting with OpenTrain AI, working remotely and contributing to real engineering work that improves how models are built and maintained.
About AI Training and Data Engineering Work
AI training includes many tasks that help models learn from human-created and human-reviewed data. This role sits at the data-engineering side of that work: building, testing and optimizing data processing systems that feed AI pipelines.
This kind of work is remote, flexible, and accessible to experienced engineers who enjoy designing reliable ETL, debugging large codebases, and improving performance at scale.
The Role
We are hiring experienced data engineers to work as Databricks Specialists. You will design and optimize large-scale data workflows, build and tune Spark jobs, and help keep Databricks pipelines performant and reliable.
This is a contractor, part-time opportunity (20+ hours/week) paid at USD $12 per hour. The role is fully remote and open worldwide; strong written English (B1 or B2) is required for collaboration and documentation.
You will build and maintain Databricks-based data pipelines and ETL, analyze and optimize Spark jobs, and debug issues in large distributed data processing systems.
Work includes hands-on coding, performance tuning, troubleshooting OutOfMemory and other stability problems, documenting fixes, and collaborating asynchronously with a remote engineering team.
Develop efficient Databricks workflows and ETL pipelines
Optimize Spark jobs for performance and memory usage
Analyze, debug, and test large codebases in Databricks
Write clear documentation and collaborate remotely
Requirements
Do not apply unless you meet the concrete, stated requirements below – we will ask for specifics during screening.
You must be able to list the exact number of years of experience you have with each programming language you know and state how many hours per week you are available for this project.
Minimum 5 years of hands-on experience working with Databricks
Deep expertise with Apache Spark, including building, optimizing, and troubleshooting Spark-based systems
5+ years of experience in at least one of: Python, Java, SQL, Scala, or Spark; indicate which language(s) and years for each
Strong experience building and optimizing data pipelines and ETL processes
Experience analyzing, debugging, and testing large code bases and navigating complex documentation
Familiarity with cloud platforms such as Azure or AWS (preferred)
Ability to work independently and solve complex technical problems
Excellent communication skills for remote collaboration
English proficiency: B1 or B2
Interview, Tests, and What We’ll Ask
The interview includes technical screening and two written test questions below. Please be ready to answer them fully; do not conclude the live interview until both are answered. When you apply, include a brief interview summary that lists the programming language(s) you know and the exact years of experience you have with each, plus your weekly availability in hours.
We will evaluate your answers for correctness and completeness and use them to score the interview.
Before interview: include each language you know and years of experience for that language, and your available hours/week
Test Question 1 — Databricks Debugging and Optimization: You are given a PySpark job in Databricks that processes a large dataset but keeps failing with an OutOfMemoryError. What steps would you take to debug and resolve this issue? Provide specific adjustments or optimizations you would apply in Da
Test Question 2 — Code Review and Documentation: Review this Python code in a Databricks notebook: data = spark.read.csv("/path/to/file.csv", header=True) filtered_data = data.filter(data["column"] > 100) result = filtered_data.groupBy("category").count() result.show() Identify issues or areas for i
How to Apply and Next Steps
Apply through OpenTrain by submitting your resume/CV and a short cover note that lists (1) every programming language you know and the exact years of experience for each, and (2) how many hours per week you can commit. Include any Databricks or Spark references you can share.
If selected, you will be invited to a live technical interview where you must answer the two test questions above. We will score technical accuracy and completeness and confirm availability before offering a contract.
Include language-by-language years of experience in your application
State your weekly availability in hours and confirm acceptance of USD $12/hr
Be prepared to answer both test questions during the live interview
Part-time contractor role building ETL pipelines and interactive dashboards using Python, Plotly/Dash, and SQL. Entry-level, remote work under 20 hrs/week at $25/hr — ideal for hands-on data analysts who document clearly and translate questions into visuals.
Join OpenTrain AI as a part-time contractor applying mathematical statistics with Python (numpy, scipy, statsmodels, pandas) to analyze messy datasets, run hypothesis tests and regressions, and produce clear, reproducible summaries; remote, <20 hrs/week at $25/hr.
Join OpenTrain AI to build end-to-end Python scraping pipelines and deliver clean structured datasets (CSV/JSON/Sheets). Remote, 20+ hrs/week at $20/hr; you must have hands-on experience with dynamic/JS-rendered sites and B2+ English.