Site Reliability Engineer Data Center
SpaceXAI · Southaven, MS; Memphis, TN ·
- Category
- SRE
- Experience
- 2+ years
You will evaluate firmware and hardware releases, investigate complex failures, manage RMA cases, and work with vendors on resolutions. You will build monitoring tools and automation, support real-time hardware troubleshooting, document reliability findings, and participate in hardware incident response and on-call rotations.
Responsibilities
- Analyze firmware packages and hardware specifications
- Run firmware security scanning and vulnerability analysis
- Identify safety issues before releases reach the data center
- Diagnose and validate complex hardware failures
- Manage RMA claims and vendor resolutions
- Troubleshoot and optimize hardware systems with operations technicians
- Develop monitoring tools, scripts, and processes
- Document failure modes, RCAs, reliability models, RMA outcomes, and hardware evaluations
- Participate in on-call rotations and hardware incident response
Requirements
- Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field, or equivalent experience
- 2+ years of hardware reliability engineering experience
- Expertise in firmware analysis, hardware specification review, and release validation
- Experience with RMA processes and vendor negotiations
- Ability to diagnose complex hardware failures
- Familiarity with data center hardware
- Python or Bash scripting proficiency
- Experience with a systems language such as C, C++, Java, or Rust
- Problem-solving and cross-functional collaboration skills