Job Title: Senior Data Engineer
Location: Washinton DC
Summary
The Senior Data Engineer / Databricks Specialist will support a federal project by developing, designing , and optimizing high-performance Spark applications and enterprise data infrastructure in AWS Databricks. The role focuses on scalable ETL/ELT pipelines using Python, PySpark, SQL, and Delta Lake; Medallion Architecture; and secure, reliable data solutions for mission-critical environments.
Responsibilities
- Develop and optimize scalable Spark applications in AWS Databricks using Python, PySpark, and SQL.
- Design, build, and maintain ETL/ELT pipelines using Medallion Architecture (Bronze, Silver, and Gold layers), ensuring data integrity, completeness, and lineage.
- Administer Databricks workspaces, clusters, scheduled jobs, access controls, Unity Catalog, and notebooks.
- Manage AWS services including S3, Lambda, IAM, Aurora, and Glue; provision infrastructure with Terraform.
- Optimize queries and Spark jobs; troubleshoot and debug large-scale data processing solutions.
- Integrate structured, semi-structured, and unstructured data from APIs and databases.
- Integrate security and compliance controls into DevSecOps CI/CD pipelines using Jenkins, GitLab CI, or equivalent tools, supporting FISMA/FedRAMP requirements, ATO reassessments, and POA&M resolution.
- Implement data quality checks, anomaly detection, validation rules, and automated lineage tracking.
- Collaborate with cross-functional Agile/Scrum teams in multi-contractor federal environments; provide technical leadership and production on-call support.
Required Qualifications
- At least 3 years of dedicated, hands-on Databricks development experience.
- Production experience with PySpark, Apache Spark, Spark SQL, and Delta Lake.
- Strong proficiency in Python and SQL and working knowledge of REST API-based data ingestion.
- Strong hands-on experience with AWS S3, IAM, Lambda, Glue, and Aurora.
- Production experience with Terraform infrastructure as code and CI/CD tools such as Jenkins, GitLab CI, or equivalent.
- Experience working in Agile/Scrum delivery environments.
- Bachelor's degree in Computer Science, Information Technology, Data Science, or a related engineering discipline.
- Active or previous High-Risk Public Trust (Tier 4) clearance.
Preferred Qualifications
- Hands-on experience with Unity Catalog, Delta Live Tables (DLT), and MLflow or Model Registry for AI/ML pipelines.
- Experience supporting federal government contracts and FISMA/FedRAMP environments.
Preferred Certifications
- Databricks Certified Data Engineer Associate
- Databricks Certified Data Engineer Professional
- Databricks Certified Machine Learning Associate
- AWS Certified Solutions Architect – Associate or Professional
- AWS Certified Data Analytics – Specialty