Site Reliability, Staff - 18350
Synopsys ·
- Seniority
- Staff
- Category
- SRE
- Experience
- 5+ years
We Are
Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon design, IP, simulation and analysis solutions, and design services. We partner closely with our customers across a wide range of industries to maximize their R&D capability and productivity, powering innovation today that ignites the ingenuity of tomorrow.
You Are
You have spent years keeping Linux infrastructure running under real pressure, the kind where a compute farm going down at 2 a.m. means a tape-out slips and someone loses sleep. You know the difference between a system that looks healthy in a dashboard and one that actually stays up when 10,000 simulation jobs hit it at once. You have debugged enough weird failures to know that the answer is rarely in the first log you check, and you are comfortable digging until you find it.
Working across cloud and on-prem does not rattle you. You have built things in Azure or AWS using Terraform and Ansible, not because it was trendy but because manual provisioning does not scale. You script in Bash and Python not to show off but to stop doing the same thing twice. When a designer in Singapore cannot access a license server, you do not guess, you pull tcpdump, check journalctl, trace the route, and fix it.
You care about uptime because you know the teams depending on your infrastructure are racing deadlines that matter. At Synopsys, you will support compute environments that power chip design work across the company, and what you keep running directly affects whether engineers can do their jobs.
What You'll Be Doing
Administer and maintain CentOS/ALMA Linux systems across Azure, AWS, and GCP cloud environments that support large-scale EDA compute farms and design workflows
Deploy and manage cloud infrastructure at scale using Terraform and Ansible, handling compute, storage, and networking for simulation, synthesis, and verification workloads
Monitor system health, capacity, and performance using Prometheus, Grafana, Elastic Search, or Splunk, catching issues before they affect design teams
Troubleshoot complex system, network, and application-level problems using tcpdump, netstat, journalctl, and other diagnostic tools to minimize downtime during critical tape-out windows
Manage license servers (FlexLM/FlexNet) and batch scheduling systems like LSF, SGE, Slurm, or Altair PBS that queue simulation and regression jobs
Coordinate tool version upgrades, freeware installs, and compatibility validation with CAD and EDA application support teams
Participate in rotational shift work (monthly rotation) and on-call weekend coverage, providing front-line support to internal design and verification teams
The Impact You Will Have
Keep compute infrastructure running so IC design, verification, and physical design teams can meet tape-out deadlines without infrastructure delays
Reduce incident response time and system downtime through proactive monitoring, automation, and structured root-cause analysis
Scale cloud infrastructure efficiently using Infrastructure as Code, enabling faster provisioning and more reliable deployments across global teams
Improve system reliability and performance tuning so simulation and synthesis jobs run faster and more predictably
Build runbooks and documentation that help junior admins resolve issues independently and reduce repeat escalations
Strengthen cross-team collaboration by translating complex infrastructure problems into clear updates for CAD, DevOps, security, and design engineering teams
Support the adoption of AI-powered tools and workflows by maintaining the underlying compute and cloud infrastructure they depend on
What You'll Need
5+ years of hands-on Linux systems administration experience, ideally supporting compute-intensive or EDA environments
Strong proficiency with CentOS, ALMA, or similar Linux distributions in production cloud environments
Hands-on experience with Azure (preferred), and working knowledge of AWS or GCP
Proven experience using Terraform and Ansible to deploy and manage large-scale compute and storage infrastructure
Strong scripting skills in Bash and Python for automation, troubleshooting, and workflow optimization
Experience managing batch scheduling systems such as LSF, SGE/Univa Grid Engine, Slurm, or Altair PBS
Familiarity with license management tools like FlexLM or FlexNet, and monitoring/logging platforms such as Prometheus, Grafana, Elastic Search, or Splunk is a plus
Who You Are
You can troubleshoot a network connectivity issue using tcpdump, netstat, traceroute, and dig without needing a runbook in front of you
You write scripts to automate repeated tasks because doing the same manual work twice feels like a waste of time
You treat internal design teams like customers, their productivity depends on your systems staying up, and you take that seriously
You can translate a complex infrastructure failure into a two-sentence update for a VP without losing the technical nuance that matters
You stay calm during on-call incidents and work methodically through root-cause analysis instead of guessing and hoping
You are comfortable coordinating across network, security, CAD, and DevOps teams to get a problem solved, even when it crosses multiple domains
The Team You'll Be Part Of
You will be part of a dynamic and collaborative Cloud Operations team that keeps our cloud environment running smoothly, securely, and reliably. Operating 24×5, we play a critical role in ensuring service availability, proactively monitoring infrastructure, and responding swiftly to incidents. We believe in teamwork, ownership, continuous learning, and a proactive approach to problem-solving. Every team member contributes to operational excellence by embracing automation, driving improvements, and sharing knowledge. Together, we don’t just keep the cloud running—we continuously evolve, innovate, and deliver reliable services that make a real difference to the business.
Rewards and Benefits
We offer a comprehensive range of health, wellness, and financial benefits to cater to your needs. Our total rewards include both monetary and non-monetary offerings. Your recruiter will provide more details about the salary range and benefits during the hiring process.