Engineering · Senior · Work from office · As per industry standards
About the role
We are seeking a Senior DevOps Engineer who will be instrumental in building and maintaining the backbone of Rhia’s technology platform. In this role, you will design, implement, and optimize the cloud infrastructure and CI/CD pipelines that allow our development teams to rapidly and safely deploy our travel SaaS application. You will ensure that our systems are highly available, secure, and can scale as our user base grows globally. The ideal candidate is passionate about automation, relentless about system stability, and excited about the opportunity to shape DevOps culture from the ground up in a company that straddles travel and AI technology.
Responsibilities
- Architect and manage our cloud infrastructure (likely on AWS/Azure/GCP) to support Rhia’s applications and data pipelines.
- Set up and maintain scalable server architectures (compute clusters, container orchestration with Kubernetes or similar, serverless components where appropriate) to handle our travel platform’s growing load.
- Design, implement, manage, and scale new and existing cloud infrastructure.
- Leverage Infrastructure as Code (IaC) tools like Terraform to manage cloud resources.
- Manage Cloud cost and review & optimize on a monthly basis.
- Respond to and remediate issues and incidents as part of an on-call rotation.
- Develop and maintain continuous integration and continuous deployment pipelines that enable frequent, reliable releases of new features.
- Automate build, test, and deployment processes, ensuring that code can move from development to production with minimal friction and maximum quality.
- Implement comprehensive monitoring, logging, and alerting solutions (using tools like Prometheus, Grafana, ELK/EFK stack, or cloud-native monitoring services).
- Keep a close eye on system performance, uptime, and anomalies.
- Set up alerts and incident response processes to address issues proactively.
- Work closely with the security best practices to harden our infrastructure.
- Manage secrets, certificates, and ensure secure configurations of servers and services.
- Implement access controls and auditing, and stay vigilant about protecting user data and maintaining compliance with relevant standards.
- Write scripts and use Infrastructure-as-Code (IaC) tools (like Terraform, CloudFormation, or Ansible) to automate provisioning and configuration of environments.
- Strive for a “cattle, not pets” approach to servers – enabling quick, reproducible deployments and recoveries.
- Use AI tools or smart automation where possible (for example, automated scaling policies based on predictive algorithms, or AI ops tools that help analyze log patterns).
- Work hand-in-hand with software engineers, QA, and data scientists to ensure that the infrastructure meets their needs.
- Advise on deployment strategies, optimize environment configurations for new features (like ensuring a new AI microservice has the right resources), and assist developers in debugging environment-specific issues.
- Regularly review system performance (response times, throughput) and identify bottlenecks.
- Optimize our use of cloud resources for both performance and cost-effectiveness – for example, using CDN for content delivery, right-sizing instances, or leveraging spot instances where applicable.
- Ensure our platform can handle peak loads (like seasonal travel surges) gracefully.
Requirements
- 5+ years of experience in DevOps, Site Reliability Engineering, or system administration with a focus on automation.
- Proven experience managing cloud infrastructure for a SaaS application is essential.
- Strong expertise with at least one major cloud provider (AWS, GCP, or Azure).
- Hands-on experience with containerization and orchestration (deploying and managing Docker containers, and orchestrating them with Kubernetes or similar systems in production).
- In-depth knowledge of CI/CD tools (Jenkins, GitLab CI, CircleCI, etc.) and version control systems (Git).
- Ability to design pipelines that include automated testing and deployment.
- Significant experience with scripting (Bash, Python, or PowerShell) and using Infrastructure as Code (Terraform, CloudFormation, or Ansible/Puppet/Chef) to automate environment setup.
- Proficiency in setting up monitoring/alerting (CloudWatch, Datadog, New Relic, etc.) and diagnosing complex system issues.
- A knack for performance tuning at various layers – application, database, network.
- Experience with log management and analysis.
- Solid understanding of network and application security in a cloud environment.
- Familiar with implementing security best practices (VPC configuration, IAM roles, security groups, encryption of data at rest/in transit).
- Experience with SaaS products and comfort in the travel/technology sector is essential.
- Excellent problem-solving skills and the ability to work under pressure during incidents.
- Good communication skills to document processes and train developers on DevOps practices.
- A collaborative attitude – eager to work with team members across the org to achieve common goals.
Nice to have
- Startup experience, indicating you can handle fast-paced and evolving requirements.
- Knowledge of compliance regimes relevant to SaaS (like SOC2, GDPR).
- Experience in systems handling travel data or transaction flows (e.g., booking systems).
- Exposure to deploying or maintaining infrastructure for AI/ML models (such as model serving frameworks, GPU instances, or big data processing like Spark).
- Bachelor’s degree in Computer Science, Information Systems, or related field (or equivalent work experience).
- Relevant certifications (AWS Certified DevOps Engineer, Certified Kubernetes Administrator, etc.).
- Familiarity with our specific tech stack (e.g., Node.js, Python).
- Experience with database administration (for databases we use, e.g., MySQL, PostgreSQL, MongoDB) and caching systems (Redis, CDN configuration).
- Experience implementing DevSecOps practices such as automated security scanning (SAST/DAST in the CI pipeline) and infrastructure testing.
- Knowledge of chaos engineering or resilience testing tools to ensure system robustness.
- Direct experience scaling a system from a small user base to a significantly larger one.
- Active participation in DevOps communities or open-source projects.
- A passion for travel or personal experience using travel apps extensively.
- Having worked in a global team or supporting systems with international users.
About RHIA
Our hotels already have great systems. They just don't talk to each other. RHIA unifies every data stream across your operation—PMS, CRM, booking engines, F&B, and beyond. One platform. Complete visibility.RHIA - "Real time Hotel Intelligence and Analytics" founded by industry veterans, hotel owners and institutional investors is looking for charismatic, bright team members to join us in shaping the next chapter of the hospitality industry.
Hospitality Technology 1-10 Est. 2025 Dubai, UAE Website