Senior backend engineer building and scaling async Python (FastAPI) REST/WebSocket APIs and the AI plumbing behind them: multi-agent LLM workflows, provider-agnostic LLM integrations, and RAG pipelines over vector stores like pgvector/Chroma. Owns data modeling, observability (p95 latency, tracing), and full CI/CD releases, remote within Brazil.
Design and scale async REST/WebSocket APIs in Python (FastAPI or similar), using dependency injection, type hints, and clean, vertical-slice architecture
Implement multi-agent workflows (sequential, concurrent, and handoff patterns) to route work among specialized LLM agents
Integrate multiple LLM providers behind a provider-agnostic layer, enabling A/B testing and cost-aware routing rather than hard-coding a single vendor
Build and maintain Retrieval-Augmented Generation (RAG) pipelines against vector stores such as pgvector, Chroma, or a managed vector search service
Design schemas and evolve data models (SQLAlchemy/SQLModel-style ORM plus migrations) that support both traditional application logic and AI-driven features
Instrument services with structured logs, distributed tracing, and cost/latency metrics, holding a real bar for p95 response time under production load
Maintain end-to-end CI/CD ownership, lint, type-check, test, package, and deploy, so releases are boring and safe
Champion AI across the delivery lifecycle, not just code generation, and set the guardrails the team works inside: how generated code gets reviewed, how provenance and licensing are checked, and what customer data may enter a prompt
Move ideas from concept to production fast, with real ownership over architecture and outcomes, not just tickets