Requirements
- Strong hands-on experience with Data-bricks and Apache Spark (PySpark/Scala).
- Experience in SQL and data transformation techniques.
- Knowledge of ETL tools and data pipeline development.
- Experience working with cloud platforms (Azure/AWS/GCP)
- Strong Azure cloud background
- Understanding of data warehousing concepts.
- Strong problem-solving and analytical skills.
- Hands-on experience with Azure Data bricks or Delta Lake, in building ETL pipelines : batch (autoloader) and Spark structured streaming
- Knowledge of data modelling and performance tuning in Spark.
- Exposure to CI/CD pipelines and DevOps practices.
- Familiarity with data governance and security practices.
- Strong hands on working experience of Unity catalog:
- Hands on exposure to Creating end to end environments : creating catalogs, schemas, tables . materialized views, functions, volumes
- Experience in building SCD 1 and SCD2 (slowly changing dimensions ) on dimension tables
- Experience in building CDC (change data capture pipelines)
- Strong hands on experience with Lakehouse federation , creating foreign catalogs to get data from external sources
- Strong understanding of databricks partitioning , Liquid clustering
Skills
- Data-bricks
- Apache Spark
- PySpark
- Scala
- SQL
- ETL
- Data Pipeline Development
- Azure
- AWS
- GCP
- Data Warehousing
- Problem-solving
- Analytical Skills
- Delta Lake
- Autoloader
- Spark Structured Streaming
- Data Modelling
- Performance Tuning
- CI/CD
- DevOps
- Data Governance
- Data Security
- Unity Catalog
- Lakehouse Federation
- Databricks Partitioning
- Liquid Clustering
