Senior Site Reliability Expert
Upserve · Remote - U.S. ·
- Work mode
- Remote
- Seniority
- Senior
- Category
- SRE
Upserve · Remote - U.S. ·
Role Summary
Our SRE team is responsible for the design, operation and reliability of Upserve’s product infrastructure. We collaborate with teams across the company to make this happen: Developers, QA, PMs, etc.
Key Responsibilities
● Initiate and contribute to continuous improvement of our software delivery processes and practices in a multi-location, multidisciplinary team to empower and accelerate product development
● Use automation extensively to design, configure, manage, and monitor systems in support of our product development teams
● Design and architect operational solutions with the specific goal of increasing the standardization, automation, repeatability, cost-efficiency and consistency of operational tasks
● Working with developers and other SREs to design and build scalable, reliable and cost-efficient Cloud infrastructure
● Adhere to and advocate for best practices, including Infrastructure as Code, monitoring, high availability, disaster recovery, security, and SRE/DevOps methodologies
● Provide timely assistance and remediation solutions during critical situations and production incidents to help resolve service problems (You will be on call for periods of time)
Required Qualifications
● Strong knowledge of Amazon Web Services
● Strong experience with Docker, Kubernetes & Linux Systems
● Experience with configuration management tools such as Chef, Puppet, Ansible, Salt
● Experience with Infrastructure as code practices: we use Terraform & OpenTofu
● Ability to read & write complex scripts using Shell
● Ability to read & understand programming languages: Python, Ruby, Go, etc.
● Good understanding of Agile development and continuous delivery best practices, software engineering tools, processes, methods and testing
● Ability to collaborate effectively with other teams
● Ability to plan, organize, prioritize and stay focused
● Good experience provisioning and managing infrastructures with high availability constraints
● Good experience with cloud cost optimization
First 90 Days: Success Outcomes
● You are a problem solver who does not shy away from tackling complexity and critical thinking
● You have a strong will to learn, grow and get out of your comfort zone
● You have great energy and passion for technology
● You are able to express yourself flawlessly in English
● You have strong interpersonal skills
Opportunity
● Lots of autonomy, flexible work culture and possibility of remote work
● Development of high traffic products, used at the global scale
● Exposure to modern and proven technology
● Opportunity to learn and expand your skill set
● Tons of growth opportunities into technical or people management roles
● Opportunity to join a fast-paced, high-growth company