Data engineer designing, building, and optimizing big data ETL pipelines (ingest, cleansing, transformation, loading, display) and ad-hoc pipelines to power AI/ML analytics, with light Linux server admin. Core stack: Python (FastAPI, async), SQL, Apache Airflow on Linux. Requires an ACTIVE Top Secret/SCI clearance with polygraph.
PLEASE NOTE: This position requires an ACTIVE Top Secret/SCI Clearance with Polygraph. To be considered for this position, you MUST have an ACTIVE Clearance Level of Top Secret/SCI with Polygraph
Position Code: 04-DM0828-1
Summary
The Data Engineer will leverage their development skills and experience, to support the successful designing, ingesting, cleansing, transformation, loading, and display of significant amounts of data.
Duties, Tasks & Responsibilities
Designing, implementing, and optimizing data extraction, cleansing, transformation, loading, replication/distribution, and large-scale ingest systems in a Big Data environment
Develop ad-hoc ETL pipelines to help analyze new data sources
Leverage Big Data holdings to power prototypical AI/ML based analytics
Developing custom solutions/code to ingest and exploit new and existing data sources
Developing data profiling, deduping logic, and matching logic for analysis
Organizing and maintaining Data Layer documentation, so others are able to understand and use it. Also, work closely with data scientists to craft data pipelines which serve the development of modern AI/ML workflows
Collaborating with teammates, other service providers, vendors, and users to develop new and more efficient methods
Light admin responsibilities for small number of Linux servers
Effectively articulating the risks and constraints associated with software solutions, based on environment
Required Experience, Skills, & Technologies
High School Diploma/GED with 10+ years of relevant software development/programming experience.
Demonstrated data analysis, parsing, and programming language experience (e.g. Python, Java) coupled with significant SQL/database experience.
Experience with the full data lifecycle, from ingest through display, in a Big Data environment.
Hands-on experience with Python-related technologies, such as environment management, async development, and FastAPI. General experience with RESTful APIs
Experience with data pipelining systems (e.g. Apache Airflow) and developing/performing ETL tasks in a Linux environment.
Desired Experience, Skills & Technologies
Experience deploying systems that leverage AI/ML technology