Senior Machine Learning Site Reliability Engineer
Helloprima · Milano, Lombardia ·
- Work mode
- Onsite
- Seniority
- Senior
- Category
- ML ai
automationawscloudcloud-nativedatadoginfrastructure-as-codekubernetesmachine-learningmlopsobservabilitypulumipysparkpythonterraform
Prima is seeking a Senior Machine Learning Site Reliability Engineer to join the Infrastructure team in Milan. You will design, build, and operate reliable, scalable systems and drive SLOs/SLIs while collaborating with software engineers on reliability improvements.
Ideal candidates have hands-on AWS, Kubernetes, Python and PySpark experience, plus familiarity with MLOps, observability with Datadog, and IaC tools like Pulumi or Terraform.
Design, build, and operate reliable and scalable systems with SLOs/SLIs. Develop automation for infrastructure and incident response; lead RCAs. Analyze performance and cost; support security best practices and threat mitigation. Proven SRE in production with cloud-native stack. Strong AWS expertise and Kubernetes experience. Solid Python and PySpark programming skills. Familiarity with MLOps and end-to-end deployment lifecycles. Private healthcare Gym discounts Wellbeing programs Mental health support