Senior Data Quality Engineer
robusta · Abu Dhabi, UAE ·
- Seniority
- Senior
- Category
- Industrial engineering
- Experience
- 3+ years
robusta · Abu Dhabi, UAE ·
We are seeking an experienced senior Databricks data quality engineer to lead the design, implementation, and automation of enterprise-scale data quality frameworks within a Databricks environment. The successful candidate will play a key role in establishing data quality controls, profiling frameworks, remediation processes, and AI-assisted quality monitoring across a large-scale data platform consisting of 170+ datasets and over 1,300 critical data elements (CDEs).
This role requires strong expertise in Databricks, PySpark, Delta Lake, MLflow, and modern data quality management practices.
Data platform & Databricks configuration
Configure and manage Databricks workspaces, compute clusters, PySpark notebooks, Delta Lake architecture, and Unity Catalog integrations. Design scalable data quality processing frameworks across 170+ datasets and 1,346 prioritized critical data elements (CDEs).
Data profiling & quality assessment
Develop AI-assisted profiling notebooks using PySpark to establish baseline data quality scores. Assess data quality across six key dimensions including:
Analyze null rates, duplicate records, invalid values, format violations, outliers, and schema drift.
Data quality rule framework
Design and build a scalable data quality rule factory using parameterized PySpark functions. Enable automated deployment of over 6,700 data quality rules without manual rule-by-rule development. Create reusable rule templates across datasets and data quality dimensions.
Pipeline quality enforcement
Integrate data quality controls within Bronze, Silver, and Gold Delta Lake layers. Implement quality gates that prevent data progression unless predefined thresholds are met. Develop reusable Databricks jobs for automated validation and monitoring.
Data cleansing & AI-driven remediation
Build automated data cleansing pipelines for:
Deploy MLflow-managed machine learning models for:
Ensure explainability of model outputs and support human-in-the-loop validation processes.
Exception management
Design failed-record handling frameworks and quarantine Delta tables. Capture failure reasons, affected CDEs, rule references, and timestamps. Develop automated reprocessing mechanisms for corrected records.
Data quality monitoring & reporting
Build Delta Lake aggregation tables for data quality metrics. Deliver data quality KPIs to Power BI dashboards including:
Configure automated alerting using Databricks SQL Alerts and Azure Monitor.
Predictive data quality analytics
Develop predictive models to identify datasets at risk of quality degradation. Support AI-assisted root cause analysis (RCA) using profiling outputs and machine learning techniques. Export and prepare remediation datasets for prioritization and governance reporting.