Site Reliability Engineer
SpaceXAI · Southaven, MS; Memphis, TN ·
- Category
- SRE
- Experience
- 5+ years
You will design monitoring architecture, improve alert quality, and provide technical leadership during severe incidents. You will run postmortems, drive corrective actions, lead cross-functional reliability projects, maintain playbooks and dependency maps, define availability objectives, and join on-call incident response rotations.
Responsibilities
- Own monitoring architecture and alert signal quality
- Use NOC feedback to improve alert suppression and redesign
- Provide technical incident leadership and bridge coordination
- Run blameless postmortems and close corrective actions
- Lead reliability projects across compute, network, storage, and facilities
- Build and maintain playbooks, runbooks, and dependency maps
- Run game days
- Define error budgets and availability objectives
- Participate in on-call rotations and SEV incident response
Requirements
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience
- 5+ years of site reliability, systems engineering, plant operations, or large-scale production operations experience
- Large-scale incident command experience
- Fleet-scale or campus-scale monitoring and observability design experience
- Experience across at least two of compute, network, storage, power, and cooling or facilities telemetry
- Experience operating playbooks or runbooks with a 24/7 operations, control room, or NOC partner
- Python or Bash scripting proficiency
- Experience with a systems language such as C, C++, Java, Go, or Rust
- Problem-solving and cross-functional collaboration skills