You will design, implement, and maintain fault-tolerant core infrastructure and performance-critical backend services. You will build APIs, serverless architectures, and secure AI workflows, including fine-tuning and inference pipelines. You will deploy, test, document, and manage releases for services from development through production.
Responsibilities
- Design, implement, and maintain fault-tolerant core infrastructure
- Architect and manage performance-critical backend services on AWS and GCP
- Implement serverless architectures and optimize cloud resource utilization
- Implement and maintain REST and WebSocket APIs
- Collaborate with cross-functional teams to define, design, and ship features
- Design and deploy scalable, reliable, and secure AI workflows
- Design and implement AI agent frameworks and multi-agent systems
- Deploy, test, and manage releases from development through production
- Deploy services using IAM, security groups, and encryption
- Write comprehensive documentation
Requirements
- 5+ years of hands-on experience developing large applications and designing end-to-end infrastructure
- Mastery of Python
- Expertise with AWS and GCP services
- Expertise in Docker and Kubernetes orchestration
- Familiarity with Terraform, CDK, or Cloud Deployment Manager
- Understanding of generative AI, LLMs, datasets, fine-tuning, and AI inference
Benefits
- Equity
- Insurance
- Flexible PTO
- WFH policy