Python/spark/ai developer
Together Staffing Group · Durban, South Africa ·
- Seniority
- Senior
- Category
- AI engineering
- Experience
- 2+ years
Together Staffing Group · Durban, South Africa ·
Our client is seeking an experienced Python / Spark / AI Developer to help replatform legacy T-SQL workloads onto modern Spark and Delta Lake pipelines while developing robust APIs and AI-enabled capabilities. This intermediate-to-senior role is suited to a developer who is comfortable working autonomously in a specification-driven environment and delivering production-quality, well-tested code.
Key Responsibilities Build production-grade Spark/Py Spark pipelines using Delta Lake. Replatform legacy T-SQL logic into Spark SQL while preserving functional parity. Develop modern, type-safe Python using strong testing and static-analysis practices. Build REST APIs using Fast API or equivalent frameworks. Implement authentication, idempotency and reliable job-status semantics. Validate migrated pipelines against legacy systems and provide evidence of output parity. Work from design documents and Architecture Decision Records (ADRs). Contribute tests and documentation alongside code changes. Participate in Git-based, PR-driven development and CI processes. Essential Requirements 4+ years' professional Python development experience. 2+ years' production experience building Spark/Py Spark or comparable distributed data pipelines. Strong Python 3.12 experience, including type hints, Pydantic, modern packaging and static analysis. Strong Apache Spark/Py Spark knowledge, including Data Frame API, Spark SQL, partitioning and performance tuning. Understanding of Spark driver/executor architecture and Spark Connect. Production experience with Delta Lake or equivalent lakehouse technologies such as Iceberg or Hudi. Strong SQL skills and the ability to translate legacy T-SQL logic into Spark SQL. Experience with MERGE operations, schema evolution, idempotent writes, hashing, surrogate keys and deduplication. Strong automated testing experience using pytest. REST API development experience, preferably with Fast API. Experience with Docker-based development environments. Strong Git and CI/CD discipline, including linting, type-checking and automated testing. Bachelor's degree in Computer Science, Engineering or equivalent experience. Ability to work independently from written specifications and technical requirements. Advantageous LLM integration and AI feature development. Experience with Ollama, v LLM, llama.cpp, LM Studio or hosted AI APIs. Data governance, PII protection and privacy engineering. Experience with Hive Metastore, Trino or Apache Ranger. Azure data technologies including ADLS Gen2, Synapse/ADF, Key Vault or Service Bus. Observability tools such as Open Telemetry, Open Lineage or structured logging. Data migration and output-parity experience. Databricks Certified Developer for Apache Spark or equivalent certification.