Engineering lead for data ingestion at Reflection: build and mentor a team of data ingestion engineers while staying hands-on, leading web crawling, ingestion pipelines, and data-lake delivery at petabyte scale. Core stack includes Ray, Beam, Spark, Airflow, Prefect, Parquet, JSONL, and WARC, with work tied to LLM training impact.
You will lead, mentor, and grow data ingestion engineers while remaining hands-on with the technical stack. You will prioritize work, run acquisition campaigns, guide architecture across crawling, pipelines, and data lakes, and connect ingestion decisions to measurable model impact.
Responsibilities
Build, mentor, and grow a team of data ingestion engineers
Lead web crawling, ingestion pipelines, and data lake delivery
Make targeted individual-contributor technical contributions
Prioritize team work and run data acquisition campaigns
Guide technical and architectural decisions
Run experiments on crawling strategies, extraction methods, and ingestion tradeoffs
Partner with research, data quality, vendors, and legal stakeholders
Onboard new sources within legal, licensing, and robots.txt constraints
Requirements
Experience building, mentoring, and growing data or infrastructure engineering teams
Experience building web-scale data acquisition or ingestion systems
Experience owning production pipelines at multi-terabyte to petabyte scale
Strong coding ability
Expertise in web crawling, data ingestion pipelines, or data lakes
Experience with Ray, Beam, Spark, Airflow, Prefect, Parquet, JSONL, WARC, and data-lake architectures
Familiarity with LLM training and evaluation
Ability to guide strategy and execution across concurrent data campaigns
Benefits
Stock options
Medical, dental, vision, and life insurance
Annual wellness allowance
Daily office lunch and dinner
22 weeks of paid parental leave
Unlimited paid time off in the U.S.
30 vacation days in the U.K.
Visa sponsorship support
Regular off-sites, happy hours, and team celebrations