Data Engineer for On-Chain Pipelines
CLR3 · Toronto, Canada ·
- Employment
- Part time
- Category
- Data engineering
Designs and operates large-scale data pipelines that decode years of Solana and Hyperliquid on-chain history into validated Parquet files for researchers, ensuring correctness via validation, checksums, and lineage docs. Core stack is SQL with Python, Rust, or Go, plus Parquet and DuckDB.
Join datastore in building reliable data pipelines that decode years of on-chain history into trustworthy Parquet files for researchers. Elevate data correctness and infrastructure while ensuring high quality and accuracy.
As a Data Engineer, you will design and run large-scale decoding pipelines, focusing on Solana and Hyperliquid history. This role demands significant experience in building production data pipelines, emphasizing SQL alongside either Python, Rust, or Go. You'll ensure correctness through rigorous validation, checksums, and lineage documentation, addressing customer inquiries about schemas and coverage.
Key Responsibilities: • Design and implement backfill pipelines at scale • Model typed schemas for various protocols • Build validation systems for data correctness • Maintain accurate versioning and documentation • Optimize storage layout for fast query performance
Requirements: • Proven experience in production data pipeline development • Proficiency in SQL plus Python, Rust, or Go • Familiarity with Parquet and tools like DuckDB • Strong focus on data correctness and validation • Willingness to work part-time in Toronto office
Help create a reliable data pipeline that researchers can trust, ensuring high-quality deliveries every time.
As a Data Engineer, you will design and run large-scale decoding pipelines, focusing on Solana and Hyperliquid history. This role demands significant experience in building production data pipelines, emphasizing SQL alongside either Python, Rust, or Go. You'll ensure correctness through rigorous validation, checksums, and lineage documentation, addressing customer inquiries about schemas and coverage.
Key Responsibilities: • Design and implement backfill pipelines at scale • Model typed schemas for various protocols • Build validation systems for data correctness • Maintain accurate versioning and documentation • Optimize storage layout for fast query performance
Requirements: • Proven experience in production data pipeline development • Proficiency in SQL plus Python, Rust, or Go • Familiarity with Parquet and tools like DuckDB • Strong focus on data correctness and validation • Willingness to work part-time in Toronto office
Help create a reliable data pipeline that researchers can trust, ensuring high-quality deliveries every time.