Project description
Our client is advancing its in-vehicle voice assistant into an intelligent, AI-powered companion. Large-language-model capabilities (Azure OpenAI / ChatGPT) have been running in production across vehicles. The goal of the project is to develop a backend which is the cloud AI orchestration service behind this: it receives requests from the vehicle, routes them, orchestrates the LLM, tool services and agents, and returns an answer or action to the car. DXC Luxoft serves as the end-to-end delivery partner, working in a joint product team with the client's engineers on the Azure platform.
This is the series development and operations work package - a live platform serving a large vehicle fleet, which extends sprint by sprint the backend features while availability and backward compatibility are maintained. New capability in the pipeline includes streaming across the full ASR → LLM → TTS chain, barge-in, multi-intent handling, a guardrails layer for deterministic vehicle-safe answers, agent routing and new tool integrations.
The role is Technical Lead and Location Lead for the engineering team. It is a hands-on delivery leadership role: the person is accountable for what the team ships, for the service running inside its availability and incident targets, and for the technical growth of the location.
Responsibilities
- Lead the backend engineering team (backend, DevOps and AI Ops engineers): technical direction, work breakdown, estimation, sprint commitment and delivery against a continuous feature and maintenance flow.
- Stay hands-on — a meaningful share of the week in code and in review. This role writes the hard parts, not just the tickets about them.
- Own engineering quality standards and their enforcement: design and code review, Python toolchain and quality gates (pytest, testcontainers, ruff, black, pyright, pydantic), test coverage expectations, definition of done.
- Implement the architecture, and push back on it when it won't survive contact with production: co-author ADRs with the Solution Architect, surface feasibility and operability concerns early, and own the implementation path once a decision is made.
- Deliver the work package roadmap: streaming migration of AI service calls (incremental chat completion, real-time ASR, TTS during synthesis) with buffering, connection management and cancellation; barge-in; multi-intent handling; guardrails and system-prompt enforcement; agent routing; new tool and agent integrations.
- Own release and version management: branch strategy, CI/CD pipeline health on Azure DevOps (incl. self-hosted runners) and GitHub, automated test gates, staged rollout, deployment and rollback.
- Be accountable for the service in production: monitoring, alerting and tracing (Azure Monitor, LangFuse, OpenTelemetry, PagerDuty routing), incident command and resolution inside the 24-hour target, root-cause analysis, stability and performance measures, patch management, and cost/latency optimisation.
- Run 2nd- and 3rd-level support processes for the location during business hours, including the on-call rota, escalation paths and handover discipline across time zones.
- Support vehicle integration and end-to-end testing together with the E2E testing work package, and participate in defect triage across the vehicle/backend boundary.
- Build and grow the location: participate in hiring and technical interviews, onboard new engineers, develop skills across the team, and keep the location's engineering reputation with the client intact.
- Represent the location in sprint ceremonies, technical alignment and client-facing reviews; maintain technical documentation in Confluence and work within the client's requirements management tooling.
SKILLS
Must have
- 8+ years in backend engineering with Python, including 2+ years leading a team of 5–10 engineers as technical lead — with the team's delivery, not just their own, as the measure.
- Still hands-on: expert Python (FastAPI, asyncio), and comfortable owning the hardest implementation in the sprint.
- Experience with streaming architectures in production (SSE/WebSocket/gRPC streaming, buffering, backpressure, cancellation semantics).
- REST and gRPC API development and operation, with OAuth2 and mTLS for secure communication and authorisation.
- PostgreSQL at depth: schema design, tuning, migrations via Alembic; plus practical handling of embeddings for RAG-based LLM queries.
- Hands-on experience integrating LLM APIs into production backends — Azure OpenAI or OpenAI API, LangChain/LangGraph or equivalent, tool/agent workflow orchestration, prompt handling and LLM constraint handling.
- Container orchestration: Docker and Kubernetes (deployment, scaling, secrets/config, networking, load balancing).
- Azure cloud experience (AKS or Container Apps, Managed Identity, Key Vault, Monitor) and Terraform/IaC: reusable modules, environment separation, remote state and locking, CI/CD integration, infrastructure lifecycle management.
- Demonstrable accountability for a production service: monitoring and alert design, distributed tracing with OpenTelemetry, incident management under a resolution SLA/SLO, patch and release management, and the judgement calls that come with being the escalation point.
- CI/CD ownership on Azure DevOps (Repos, Pipelines) and/or GitHub, with automated testing, linting and formatting gates.
- Demonstrable Scrum experience (sprints, reviews, retros, backlog maintenance) and the discipline to keep estimates honest.
- English C1 — daily technical communication with the client's engineers and the onshore team.
- Working-hours overlap with CET sufficient for daily alignment, and willingness to travel to Germany for knowledge transfer, workshops and integration phases — including an intensive onboarding/takeover period at project start.
Nice to have
• Experience taking over a live system from a previous supplier: knowledge transfer under time pressure, reverse-documenting an inherited codebase, stabilising what you did not build.
• Experience building up an offshore location or team from a small core — hiring, onboarding, retention, and establishing credibility with a demanding customer.
• Automotive or in-vehicle software context: MIB3 / E³ / SDV platforms, Android Automotive (AAOS), AIDL, Viwi protocol, vehicle telemetry and configuration services.
• Streaming-level ASR/TTS service integration (Azure Speech, Cerence, Google TTS, Whisper).
• LLM observability and evaluation practice: LangFuse, prompt regression suites, quality metrics (faithfulness, hallucination rate, latency P95, WER/CER).
• Broader persistence experience: MongoDB (pymongo/beanie), CosmosDB.
• Awareness of automotive quality and security frameworks (A-SPICE, ISO/SAE 21434, TISAX) and of privacy-by-design/GDPR constraints on telemetry and logging.
• German language skills.