DevOps / Site Reliability Engineer
IMG Academy · US ·
- Work mode
- Remote
- Category
- Devops
- Experience
- 3+ years
IMG Academy · US ·
Workwize · Amsterdam
Restroworks · Delhi, India
Next College Student Athlete (NCSA) · United States
AvePoint · Singapore
Position Summary:
We’re looking for a DevOps / Site Reliability Engineer to join the team supporting SportsRecruits, our sports recruiting platform connecting student-athletes, clubs, events and college coaches throughout the recruiting process. You will play a key role in ensuring our systems are efficient, reliable, and scalable, while helping us improve developer productivity and application performance.
You’ll collaborate closely with developers, QA, product, and our cloud security engineer to streamline builds and deployments, maintain application infrastructure, and proactively solve issues before they impact our users.
We are a product development team of committed engineers, designers, and product managers distributed across the United States. We are scaling the SportsRecruits network and building innovative tools to empower student-athletes, college coaches, and event operators. Our tools are built on technologies spanning mobile and web applications, computer vision, and LLMs.
We emphasize performance, security, and maintainability—and we love solving problems that have real-world impact on student-athletes, coaches, and partners.
Position Responsibilities:
CI/CD & Deployments:
Configure, manage, and improve Bitbucket pipelines for deploying our applications to staging and production.
Improve CI pipeline speed, reliability, and security in collaboration with our Cloud Security Engineer.
Assist developers and QA teams with deployments.
Work with Docker and AWS ECR for container builds and deployment workflows.
Monitoring & Incident Response:
Review and investigate system issues flagged by Sentry, NewRelic, and CloudWatch.
Monitor application performance, identify bottlenecks, and propose solutions.
Respond to production and staging issues, including database latency, unresponsive resources, or failed jobs.
Environment & Infrastructure Management:
Maintain and support non-production environments used by developers and QA.
Maintain and improve AWS infrastructure and Terraform resources.
Perform updates and upgrades to AWS services as needed to ensure reliability and ability to scale.
Collaboration & Continuous Improvement:
Partner with engineers to design systems that are scalable, observable, and resilient.
Work closely with our cloud security engineer to ensure secure configurations in CI/CD, AWS, and containerized workloads.
Contribute ideas and improvements to workflows, automation, and monitoring strategies.
Leverage AI to automate monitoring and diagnosis.
Knowledge, Skills and Abilities:
3+ years of experience in DevOps, SRE, or related engineering roles.
Strong experience configuring CI/CD pipelines (Bitbucket Pipelines, GitHub Actions, or similar).
Experience configuring, debugging and deploying PHP applications
Hands-on experience with Docker and AWS ECR for container builds and deployments.
Strong experience with AWS services (EC2, RDS, ECS, Lambda, etc.) and Terraform for infrastructure as code.
Familiarity with monitoring and observability tools such as New Relic, Sentry, CloudWatch, or similar.
Strong troubleshooting skills for debugging performance issues in databases, applications, and distributed systems.
Experience with modern software development workflows (agile teams, code reviews, branching strategies).
Strong scripting and automation skills (Bash, Python, or similar).
Excellent communication skills and a collaborative mindset.
Interest in leveraging AI agents to automate monitoring and diagnosis workflows
#LI-TR1
IMG · Uxbridge, United Kingdom