About the Role
This role involves planning and accomplishing goals, performing complicated tasks, and utilizing creativity and latitude to solve business problems. The individual will analyze user requirements, identify system issues, and design solutions to improve existing computer systems. Responsibilities include conferring with personnel to analyze procedures, identify problems, and define input/output requirements, as well as writing detailed descriptions of user needs and program functions. The role also requires reviewing system capabilities and limitations to determine the feasibility of program changes.
Responsibilities
- Design, develop, and optimize scalable data solutions on Databricks using PySpark or Scala.
- Build and maintain ingestion pipelines, Declarative Pipelines (DLT), and Medallion Architecture (Bronze, Silver, Gold).
- Develop robust data models and implement data quality, validation, and governance frameworks.
- Create dynamic dashboards, Databricks Apps, and analytical solutions.
- Optimize workloads, monitoring, and operational processes for scalability, security, and cost efficiency.
- Analyze user requirements, procedures, and problems to automate processing or improve existing computer systems.
- Confer with personnel to analyze current operational procedures, identify problems, and learn specific input and output requirements.
- Write detailed descriptions of user needs, program functions, and steps required to develop or modify computer programs.
- Review computer system capabilities, specifications, and scheduling limitations.
Requirements
- 8 or more years of experience in IT, supporting the design, development, deployment, or delivery of technology solutions.
- 8 or more years of experience with Databricks, including building and optimizing ETL/ELT data pipelines using Apache Spark.
- 8 or more years of experience in data warehousing and dimensional data modeling (star/snowflake schemas).
- 8 or more years of experience implementing data governance, data quality, and data security practices.
- 8 or more years of experience implementing Lakeflow Declarative Pipelines (formerly Delta Live Tables/DLT) for building and managing production data pipelines.
- 8 or more years of experience with Delta Lake, medallion architecture (bronze/silver/gold layers), data lakehouse design, and creating and scheduling offline jobs using Lakeflow Jobs (formerly Databricks Workflows) or similar orchestration tools (e.g., Airflow).
- 1 year of experience working in public sector or state government environments (Preferred).
- Experience with CI/CD practices for data pipelines (DevOps, Git-based workflows) (Preferred).
Skills
- PySpark
- Scala
- SQL
- Python
- Databricks
- Apache Spark
- ETL/ELT
- Data Warehousing
- Dimensional Data Modeling
- Star Schemas
- Snowflake Schemas
- Declarative Pipelines (DLT)
- Medallion Architecture (Bronze, Silver, Gold)
- Delta Lake
- Data Lakehouse Design
- Lakeflow Jobs (Databricks Workflows)
- Airflow
- Databricks SQL dashboards
- Databricks Apps
- Data Governance
- Data Quality
- Data Security
- CI/CD practices
- DevOps
- Git-based workflows
- Communication skills (verbal and written)
- Presenting insights to technical and business stakeholders
Experience Level
- 8 or more years of experience
- 1 year of experience (Preferred)
Education Level
- Databricks certification (e.g., Databricks Certified Data Engineer Associate/Professional) (Preferred)
