Role Overview
We are seeking a highly skilled, autonomous Lead Data Engineer to architect and implement an automated
operational integration pipeline. The effort is focused on building an automated operational integration that
extracts required data from the Eclipse data share, snapshots into an internal Ops database, generates
InvestorTools-compatible files, and securely delivers those files through an SFTP process on a scheduled
basis with an initial full file followed by ongoing delta files.
The core objective of this initiative is to replace manual data movement and reconciliation activities with a
reliable, scalable, and fully supportable integration solution. As the Lead Engineer on this initiative, you must be
a self-starter capable of driving end-to-end technical execution with minimal oversight while engaging directly
with stakeholders and client teams.
PRIMARY RESPONSIBILITIES
1. Data Integration & ETL Engineering
Design and deploy robust ETL/ELT pipelines to extract datasets from the Eclipse data share, create historical
snapshots within the Ops database, and implement delta/change-data-capture (CDC) extraction logic.
2. Workflow Orchestration
Write, schedule, monitor, and troubleshoot automated DAGs using orchestrators such as Apache Airflow to guarantee
reliable, scheduled pipeline execution and file delivery.
3. Database & SQL Management
Apply advanced SQL skills to perform complex data transformations, query data-share/warehouse environments, and
manage relational schemas within the Ops database.
4. Automation & SFTP File Generation
Utilize Python to construct custom file parsing, formatting, and schema validation routines for InvestorToolscompatible file outputs and secure automated SFTP transmission.
5. Client Communication & Leadership
Act as the primary technical point of contact for client stakeholders. Communicate system architecture, progress, and technical requirements clearly while working independently as a self-starter
Qualification Matrix
Minimum Qualifications
5+ Years Experience: Deep hands-on experience
designing ETL/ELT pipelines, data snapshots, and
CDC extractions.
Orchestration Expertise: Practical proficiency
writing, scheduling, and troubleshooting DAGs in
Apache Airflow.
Advanced SQL: Strong experience in relational
schema management, warehouse queries, and
data transformation.
Python Scripting: Proficient in Python for custom
file parsing, formatting, schema validation, and
SFTP automation.
Financial Domain Exposure: Direct experience
with financial datasets (portfolio accounting,
trading, or fixed income).
Client Communication: Clear verbal and written
communication skills for direct client
engagement.
Independent Worker: Strong self-starter ability to
operate with minimal dependencies or
supervision.
Nice to Have
Platform Experience: Direct prior experience with
Eclipse data share environments or
InvestorTools software integrations.
Fixed-Income Analytics: Understanding of fixed income security processing, yield calculations,
and bond accounting rules.
Data Security: Experience with automated
encryption (PGP/SSH), secure key management,
and enterprise SFTP protocols.
CI/CD & DevOps: Exposure to automated
deployment pipelines (Docker, GitHub Actions,
Jenkins) for data workflows.
Data Quality Frameworks: Knowledge of
automated data testing frameworks (Great expectation dbt tests)
Key Operational Objectives
Manual Effort Elimination: Completely replace existing manual spreadsheet uploads and ad-hoc reconciliation
activities with automated, supportable data pipelines.
Robust Delta Processing: Ensure fault-tolerant execution of initial full data loads and continuous daily incremental
delta generations without data loss or duplication.
Stakeholder Alignment: Serve as an effective technical lead who communicates milestone status, technical risks,
and interface requirements directly to client managers and engineering peers.