Databricks architect
EXL · Pune, Maharashtra, India ·
This posting has been closed.
- Work mode
- Hybrid
- Category
- Architecture
Key Responsibilities ● · Design, develop, and maintain data pipelines using Databricks, Apache Spark, and cloud-native services. ● · Build and optimize ETL/ELT workflows for large-scale structured and unstructured data. ● · Develop data models and implement data quality, validation, and governance frameworks. ● · Integrate data from multiple sources into a unified Lakehouse architecture. ● · Optimize Spark jobs and Databricks workloads for performance, scalability, and cost efficiency. ● · Implement security controls, access management, and data governance using Unity Catalog. ● · Collaborate with business, analytics, and AI/ML teams to deliver trusted data products. ● · Monitor, troubleshoot, and resolve data pipeline issues. ● · Support CI/CD, DevOps, and infrastructure automation practices. ● · Maintain technical documentation and best practices. Required Technical Skills /Core Technologies ● · Databricks Lakehouse Platform ● · Apache Spark / PySpark ● · Delta Lake ● · SQL ● · Python ● · ETL / ELT Development ● · Data Modeling ● · Data Warehousing ● · Data Quality & Validation ● · Streaming & Real-Time Processing ● Governance & Security ● · Unity Catalog ● · Data Lineage ● · Row-Level Security ● · Access Control & Compliance ● · Data Governance Frameworks ● Cloud & DevOps ● · Azure / AWS / GCP ● · Terraform ● · GitHub Actions / Azure DevOps ● · CI/CD Pipelines ● Analytics & AI ● · Semantic Layers ● · Data Products ● · BI Platforms ● · Machine Learning Support ● · Generative AI & RAG Architectures
Key Responsibilities ● · Design, develop, and maintain data pipelines using Databricks, Apache Spark, and cloud-native services. ● · Build and optimize ETL/ELT workflows for large-scale structured and unstructured data. ● · Develop data models and implement data quality, validation, and governance frameworks. ● · Integrate data from multiple sources into a unified Lakehouse architecture. ● · Optimize Spark jobs and Databricks workloads for performance, scalability, and cost efficiency. ● · Implement security controls, access management, and data governance using Unity Catalog. ● · Collaborate with business, analytics, and AI/ML teams to deliver trusted data products. ● · Monitor, troubleshoot, and resolve data pipeline issues. ● · Support CI/CD, DevOps, and infrastructure automation practices. ● · Maintain technical documentation and best practices. Required Technical Skills /Core Technologies ● · Databricks Lakehouse Platform ● · Apache Spark / PySpark ● · Delta Lake ● · SQL ● · Python ● · ETL / ELT Development ● · Data Modeling ● · Data Warehousing ● · Data Quality & Validation ● · Streaming & Real-Time Processing ● Governance & Security ● · Unity Catalog ● · Data Lineage ● · Row-Level Security ● · Access Control & Compliance ● · Data Governance Frameworks ● Cloud & DevOps ● · Azure / AWS / GCP ● · Terraform ● · GitHub Actions / Azure DevOps ● · CI/CD Pipelines ● Analytics & AI ● · Semantic Layers ● · Data Products ● · BI Platforms ● · Machine Learning Support ● · Generative AI & RAG Architectures
Key Responsibilities ● · Design, develop, and maintain data pipelines using Databricks, Apache Spark, and cloud-native services. ● · Build and optimize ETL/ELT workflows for large-scale structured and unstructured data. ● · Develop data models and implement data quality, validation, and governance frameworks. ● · Integrate data from multiple sources into a unified Lakehouse architecture. ● · Optimize Spark jobs and Databricks workloads for performance, scalability, and cost efficiency. ● · Implement security controls, access management, and data governance using Unity Catalog. ● · Collaborate with business, analytics, and AI/ML teams to deliver trusted data products. ● · Monitor, troubleshoot, and resolve data pipeline issues. ● · Support CI/CD, DevOps, and infrastructure automation practices. ● · Maintain technical documentation and best practices. Required Technical Skills /Core Technologies ● · Databricks Lakehouse Platform ● · Apache Spark / PySpark ● · Delta Lake ● · SQL ● · Python ● · ETL / ELT Development ● · Data Modeling ● · Data Warehousing ● · Data Quality & Validation ● · Streaming & Real-Time Processing ● Governance & Security ● · Unity Catalog ● · Data Lineage ● · Row-Level Security ● · Access Control & Compliance ● · Data Governance Frameworks ● Cloud & DevOps ● · Azure / AWS / GCP ● · Terraform ● · GitHub Actions / Azure DevOps ● · CI/CD Pipelines ● Analytics & AI ● · Semantic Layers ● · Data Products ● · BI Platforms ● · Machine Learning Support ● · Generative AI & RAG Architectures