About the Role
The Databricks Engineer designs, develops, and optimizes scalable data solutions on the Databricks platform, leveraging PySpark or Scala for large-scale data processing. This role involves constructing ingestion pipelines, implementing medallion architecture, and designing dimensional data models. The engineer collaborates with cross-functional stakeholders to gather requirements, perform cost/benefit analyses, and implement data governance, security, and quality frameworks.
Responsibilities
- Design, develop, and optimize highly scalable data solutions on Databricks using Apache Spark, PySpark, or Scala for large-scale data processing.
- Build, configure, and maintain robust ETL/ELT ingestion pipelines utilizing Lakeflow Declarative Pipelines.
- Implement and maintain Delta Lake architectures and medallion database structures across Bronze, Silver, and Gold data layers.
- Orchestrate, schedule, and monitor offline production jobs utilizing Lakeflow Jobs or equivalent tools.
- Design, develop, and support comprehensive enterprise data warehouses utilizing dimensional data modeling, including Star and Snowflake schemas.
- Implement rigorous data quality validation, data security controls, platform access models, and data governance frameworks.
- Analyze system specifications, evaluate operational limitations, and perform data audits to ensure scalability, security, and cost efficiency.
- Collaborate directly with business stakeholders, program managers, and technical peers to understand operational objectives, identify problems, and analyze current procedures.
- Translate high-level business goals into formal technical requirements, system design documentation, and cost/benefit analyses.
- Develop dynamic analytical dashboards and reporting solutions natively within Databricks, including Databricks SQL dashboards and Databricks Apps, to deliver actionable operational insights.
Requirements
- 8 or more years of experience in IT, supporting the design, development, deployment, or delivery of technology solutions.
- 8 or more years of experience with Databricks, including building and optimizing ETL/ELT data pipelines using Apache Spark.
- 8 or more years of experience in data warehousing and dimensional data modeling, including Star and Snowflake schemas.
- 8 or more years of professional proficiency utilizing SQL and Python (or Scala) for large-scale data processing.
- 8 or more years of experience designing and developing dashboards and applications natively within Databricks, such as Databricks SQL dashboards or Databricks Apps.
- 8 or more years of experience implementing data governance, data quality metrics, and data security practices.
- 8 or more years of experience implementing Lakeflow Declarative Pipelines to build and manage production pipelines.
- 8 or more years of experience with Delta Lake, medallion architecture, data lakehouse design, and scheduling offline jobs using Lakeflow Jobs or similar orchestration tools.
- 8 or more years of experience communicating technical specifications and presenting data-driven insights to technical and non-technical stakeholders.
- 1 or more years of experience working within public sector or state government environments.
- 1 or more years of experience implementing CI/CD practices for data pipelines, including DevOps and Git-based version control workflows.
Skills
- Apache Spark
- PySpark
- Scala
- ETL/ELT
- Lakeflow Declarative Pipelines
- Delta Live Tables
- Delta Lake
- Medallion architecture
- Bronze data layer
- Silver data layer
- Gold data layer
- Lakeflow Jobs
- Databricks Workflows
- Apache Airflow
- Dimensional data modeling
- Star schema
- Snowflake schema
- Data quality validation
- Data security controls
- Platform access models
- Data governance
- SQL
- Python
- Databricks SQL dashboards
- Databricks Apps
- CI/CD practices
- DevOps
- Git-based version control
- Analytical skills
- Critical-thinking skills
- Problem-solving skills
- Communication skills
Location
- Austin, TX
Work Type
- Remote
- Flexible work from home options available
Experience Level
- Senior Level
- 8 or more years of experience
Education Level
- Databricks Certified Data Engineer Associate or Professional
Benefits
- Competitive salary
