Data engineer role in Singapore centered on building and tuning analytical platforms around Apache Doris, Spark, and Iceberg. Day to day involves writing and optimizing complex SQL, managing lakehouse table formats and MPP query execution, and automating data pipelines with Python/Java/Scala plus Git and CI/CD.
Job Title:
Data Engineer - Apache Doris
Data Engineer -
• Hands-on production experience with Apache Doris (or a comparable MPP OLAP engine - StarRocks, ClickHouse, Greenplum).
• Strong, deep SQL expertise - complex analytical queries, window functions, CTEs, query optimization.
• Solid understanding of Doris architecture (FE/BE), the three data models
(Duplicate/Aggregate/Unique), partitioning, bucketing, and tablet/replica management.
• Experience with Apache Iceberg (or comparable open table formats - Delta Lake, Hudi) and lakehouse / external-catalog federation.
• Practical experience with Snowflake and/or Apache Spark- enough to read, understand, and migrate existing workloads.
• Understanding of distributed / MPP query execution: join distribution, runtime filters, memory management, and data skew.
• Hands-on production experience with Trino (or PrestoSQL/Presto)
• Understanding of distributed query execution: MPP architecture, join distribution, memory/spill behaviour, partition pruning, and predicate pushdown
• Experience with cloud object storage and columnar file formats (Parquet, ORC)
• Proficiency in at least one programming language (Python, Java, or Scala) for tooling, UDFs, and Automation
• Version control (Git) and CI/CD for data pipelines.
Mandatory skills*
Apache Doris / Spark/Iceberg + SQL
Desired skills*
Apache Iceberg / Lakehouse technologies, Advanced SQL, Trino/Presto, Snowflake, Spark