We are looking for a Data Engineer with deep expertise in Databricks to design, build, and scale our data platform. You will own the architecture and performance of large-scale data pipelines, drive best practices across the Lakehouse stack, and partner closely with Data Science, Analytics, and AI Engineering teams to deliver reliable, production-grade data infrastructure.
This role is ideal for someone who has moved beyond writing pipelines to also shaping platform standards, mentoring engineers, and making architectural decisions on a Databricks-based Lakehouse.
Key Responsibilities
- Design, build, and optimize scalable ETL/ELT pipelines on Databricks, processing large volumes of structured and unstructured data
- Architect and maintain Delta Lake-based data models, ensuring data quality, reliability, and performance across Bronze/Silver/Gold layers
- Develop and orchestrate workflows using Databricks Workflows, Delta Live Tables (DLT), and job scheduling tools
- Implement data governance, access control, and lineage using Unity Catalog
- Optimize Spark jobs for performance and cost (cluster sizing, partitioning, caching, Photon, Z-ordering, file compaction)
- Build and maintain CI/CD pipelines for Databricks notebooks/jobs (Databricks Asset Bundles, Repos, Git integration)
- Collaborate with Data Scientists and ML Engineers to support feature engineering and MLflow-based model lifecycle workflows
- Define and enforce data engineering best practices: testing, version control, observability, and documentation
- Monitor pipeline health and cost, implementing alerting and proactive optimization (cluster policies, job clusters vs. all-purpose clusters)
- Mentor mid-level engineers and contribute to technical decision-making and platform roadmap
- Ensure compliance with data security and privacy requirements (e.g., GDPR)
Must-Have Skills
- 5+ years of experience in Data Engineering, with 3+ years hands-on with Databricks
- Databricks Technical Administration and Maintenance (SME-level: knowledge acquisition, ongoing support and maintenance of the Databricks platform)
- Strong proficiency in Apache Spark (PySpark and/or Scala)
- Solid experience with Delta Lake architecture (Bronze/Silver/Gold, ACID transactions, schema evolution)
- Hands-on experience with Unity Catalog for governance and access management
- Strong SQL skills and experience with Databricks SQL Warehouses
- Experience building and orchestrating pipelines with Delta Live Tables and/or Databricks Workflows
- Proficiency in Python
- Experience with CI/CD for data pipelines (Databricks Asset Bundles, Git, Azure DevOps/GitHub Actions/GitLab CI)
- Strong understanding of cloud data architecture (any of Azure, AWS, or GCP)
- Experience with performance tuning and cost optimization on Databricks clusters
Nice-to-Have Skills
- Cloud-specific Databricks experience: Azure Databricks, Databricks on AWS, or Databricks on GCP
- Experience with MLflow for experiment tracking and model registry
- Familiarity with Terraform for infrastructure-as-code (Databricks provider)
- Exposure to streaming data pipelines (Structured Streaming, Kafka, Event Hubs/Kinesis)
- Experience with dbt on Databricks
- Knowledge of data observability tools (e.g., Monte Carlo, Great Expectations)
- Familiarity with LLMOps/AI workloads on Databricks (Vector Search, Model Serving)
- Databricks certifications (Data Engineer Professional/Associate, Spark Developer)
- Experience working under compliance frameworks (GDPR, ISO 27001, DORA)