The Role:
Builds performant data processing pipelines, manages relational databases, and constructs cloud-portable data lake environments using modern open storage formats and embedded analytical
engines.
architecture principles (Bronze/Silver/Gold) to transform raw data into analytics-ready assets.
Apache Iceberg, and Delta Lake to ensure cross-cloud compatibility.
performant localized data processing within ETL workflows.
SQLite for embedded or lightweight localized storage needs.
Databricks or Apache Spark
Requirements:
Advanced SQL proficiency with deep PostgreSQL expertise and working knowledge of SQLite.
Hands-on experience building robust analytical pipelines using Python and DuckDB.
Mastery of open table and storage formats (Parquet, Apache Iceberg, Delta Lake).
Proven track record designing data lakes using the Medallion data processing architecture.
Familiarity with the Databricks / PySpark ecosystem for distributed processing.
Python & SQL | PostgreSQL & SQLite | Medallion Architecture | DuckDB | Parquet / Iceberg / Delta | | Databricks / Spark
C - PGS - 29092026
Wakapi Web