Senior Data Platform Engineer
ARUKAH CAPITAL PTE. LTD. · Singapore, Singapore ·
- Seniority
- Senior
- Category
- Data engineering
ARUKAH CAPITAL PTE. LTD. · Singapore, Singapore ·
Own Arukah's end-to-end data platform for biochar and biogas climate measurement (dMRV) — from ingesting sensor, lab, document and operational data to analytics and registry reporting. A hands-on Python/SQL role on GCP (Cloud Run, Dagster, PostgreSQL), re-architecting for scale from a few plants to many, using AI coding assistants daily.
Arukah builds technology that connects biochar and biogas operations to measurable climate impact. Our systems bring together production records, sensors, laboratory results and supporting documents for digital measurement, reporting and verification (dMRV).
Our next challenge is to evolve the technology supporting a few plants into a platform for 10s of hundreds of plants, with an elite small engineering team. We need reusable systems, trustworthy data and thoughtful choices about what to build and what to obtain from managed services.
Own the data platform from source ingestion through operational analytics and registry reporting. You will assess the existing codebase, shape the architecture and deliver improvements incrementally while supporting current operations. The responsibilities include re-architecting where needed, not simply maintaining existing pipelines.
This role combines platform engineering, data engineering and analytics engineering. We are looking for strong platform judgment and hands‑on Python and SQL skills, supported by solid analytical modeling. You should be comfortable choosing a managed capability, building a domain‑specific component, or simplifying a workflow based on reliability, correctness, cost and the team’s capacity to operate it.
You will use AI coding assistants as part of everyday engineering, with responsibility for the design, correctness and maintainability of everything delivered. You will also help operate AI‑based document extraction and human review within the data workflow.
Architecture for multiple plants. Design shared capabilities, facility configuration, access boundaries and data isolation. Establish a practical path from a few plants to many plants, using representative workloads and operational needs to guide decisions. Make each additional plant easier to onboard without creating separate code forks.
Managed‑service and build‑versus‑buy decisions. Evaluate compute, storage, databases, orchestration, connectors, identity and monitoring. Consider engineering time, service limits, recurring cost, recovery, vendor dependencies and exit options. Prefer managed services when they meet requirements and reduce ongoing operational work.
Reliable ingestion and processing. Build reusable integrations for operational sheets, sensors, APIs, documents and laboratory results. Handle duplicates, late or missing data, schema changes, retries and historical reprocessing. Make source failures visible and recoverable.
Shared data models and analytics. Model plants, equipment, production batches, feedstock deliveries, samples and evidence with explicit grain, stable identifiers, units, time semantics and history. Create tested datasets and consistent metrics for plant operations, portfolio analysis and reporting.
Traceable calculations and evidence. Preserve source provenance, human corrections, calculation versions and approval states. Implement carbon and reporting rules with domain specialists, reconcile outputs and assess the effect of changes on historical results. Distinguish estimates, submitted evidence and accepted registry outcomes.
Improve identity, authorization, service permissions and secrets management across APIs, internal tools and automated workloads. Make sensitive data access and reviewer actions attributable to the right people and services.
Operate and improve our GCP and container infrastructure, including Cloud Run services, Dagster execution, PostgreSQL connectivity, storage and caching. Establish useful monitoring, alerts, recovery procedures and cost visibility.
Make development and releases repeatable through dependency management, automated tests, CI/CD, configuration management and infrastructure automation. Provide documented deployment and rollback paths.
Create operations and delivery that a small team can sustain. Establish automated checks, CI/CD, infrastructure configuration, useful monitoring, cost visibility and recovery procedures. Deliver migrations in bounded steps and document the system so another engineer can deploy, diagnose and recover it.
Effective AI‑assisted engineering. Use coding assistants and agents for codebase exploration, implementation, refactoring, test development, documentation and debugging. Build repeatable ways to provide context, review changes and verify results while keeping ownership with the engineer.
Experience owning production data systems through design, architectural change, migration and operation. You can explain the trade‑offs and measurable outcomes of your decisions.
Strong Python and SQL, including maintainable interfaces, APIs, relational databases, warehouse transformations, numerical correctness and production debugging.
Strong data modeling: record grain, keys, temporal relationships, schema evolution, lineage, units and consisten