Job Description & Details
This is a heavy-duty data engineering gig centered entirely around modern lakehouse tech. You'll be knee-deep in Databricks, building out scalable pipelines and wrestling with performance tuning on massive datasets. If you enjoy optimizing Spark jobs and designing clean Medallion architecture, you'll actually like what you're building here.
What You'll Actually Be Doing
You'll spend your days writing robust Python and SQL to build, orchestrate, and maintain high-performance batch and streaming ETL pipelines inside Databricks. Expect to spend a significant amount of time debugging Spark execution plans, managing Delta Lake storage layers with ACID compliance, and setting up automated workflows using tools like Azure Data Factory or Databricks Asset Bundles. Collaboration is also key, so you'll be pushing code through Git and working closely with analytics teams to make sure data is clean, reliable, and query-ready.
The Core Tech Stack
You need absolute fluency in Apache Spark, Databricks, Python, and advanced SQL. The team relies heavily on Delta Lake and Medallion Architecture principles, so you can't just fake your way through lakehouse concepts—you need to know how Time Travel and ACID transactions work in practice. Experience with cloud infrastructure like Azure and workflow orchestrators like ADF or Airflow will make or break your ability to keep these pipelines humming without blowing up the cloud budget.
Interview Expectations
Expect them to throw a gnarly performance tuning scenario your way, asking you to diagnose why a specific Spark job is spilling to disk or suffering from data skew. The hiring manager isn't just looking for a textbook definition of shuffle partitions; they want to hear your methodical approach to reading DAGs, optimizing joins, and adjusting cluster configs. You'll also likely get grilled on Delta Lake internals and how you manage schema enforcement and evolution in production environments.
Application Advice
Make sure your resume doesn't just list tools, but actually highlights the scale of the data you've handled and the measurable performance gains you've achieved. Weave in exact keywords from the prompt like Delta Lake, Medallion Architecture, Spark SQL, and Databricks Asset Bundles so you clear the ATS. Emphasize your production experience with Python and Git workflows rather than just sandbox projects, as they want someone who can hit the ground running without breaking CI/CD pipelines.