Senior DevOps Engineer, AI Platform
Webhosting · Canada ·
- Seniority
- Senior
- Employment
- Full time
- Category
- Devops
Webhosting · Canada ·
Network Solutions · Argentina
Newfold Digital · Canada - Remote
Aerospike · Remote (United States)
CITA LABS · Singapore, Singapore
A hands-on Senior DevOps Engineer builds and operates production cloud infrastructure on Azure (AKS) and Oracle Cloud (OKE) for AI platforms — LLM gateways, Python agent runtimes, RAG workers — plus web apps, backend services, and shared platform tooling, using Kubernetes, Terraform, Helm, Jenkins/ArgoCD CI/CD, and full observability stacks.
At Network Solutions, we’ve been trusted for decades to help people get online and stay ahead. We’ve been here since the beginning of the internet, and we’re still building for what comes next.
As the original digital identity authority, we help secure domain names, protect brands, and safeguard the infrastructure businesses rely on. We empower our customers to own and manage the assets that define them online, while delivering enterprise-grade security to protect against virtual threats. Our team leverages modern, AI-accelerated tools to streamline how businesses manage their digital presence, making the most of our decades of experience.
The Network Solutions team is here to help online businesses protect what’s theirs and build for tomorrow. That’s why millions trust us to protect their domains, brands, and websites every day.
We are looking for a hands-on Senior DevOps Engineer to build and operate the infrastructure powering our AI platforms, agent runtimes, web applications, backend services, APIs, and shared platform capabilities. Our environment includes a centralized LLM gateway, Python-based agent runtimes, RAG workers, MCP services, asynchronous processing, databases, caches, queues, and observability services.
You will work with AI engineers, application engineers, and architects who define technical designs, then independently translate those designs into reliable, scalable, secure, and observable production infrastructure across Microsoft Azure and Oracle Cloud Infrastructure.
You will frequently receive a technical design for a new AI workload, application, backend service, or platform capability. From that design, you should be able to independently determine and implement the infrastructure needed to run it in production.
AI or machine learning infrastructure experience is helpful, but not required. Strong experience with Kubernetes, web and backend infrastructure, networking, queues, databases, CI/CD, observability, and production cloud operations is the foundation for this role.
Madiff · Warsaw, Poland