Infrastructure Engineer
BETBY ·
- Category
- Devops
- Experience
- 3+ years
ansibleautomationbashdockergrafanakafkakuberneteslinuxobservabilityopensearchprometheuspythonterraform
As a fast-growing, award-winning company, Betby powers the industry with our premium sportsbook, featuring world-class risk management and seamless omni-channel support, reaching millions of players across countless markets.
Responsibilities:
- Designing, deploying, configuring, and maintaining scalable monitoring, logging, and alerting platforms for production and beta/test/dev environments;
- Operating Prometheus, Alertmanager, Grafana, Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards, including upgrades, reliability, availability, capacity, and retention planning;
- Building and maintaining reliable metrics and log collection pipelines for infrastructure and business-critical services;
- Creating dashboards that provide clear, useful visibility into service health, performance, capacity, and operational risks;
- Designing, tuning, and maintaining actionable alert rules and notification routing; reducing alert noise and improving incident response;
- Monitoring infrastructure and application metrics and logs, troubleshooting issues, and improving stability and performance under heavy loads;
- Managing metric cardinality, log volume, retention, storage consumption, and query performance to keep observability platforms scalable and cost-effective;
- Establishing high-availability and recovery approaches for observability services and validating operational readiness;
Requirements:
- Minimum 3 years of experience with administering Linux systems and operating monitoring, logging, or observability systems;
- Experience with Debian-based systems;
- Experience with Docker and Kubernetes;
- Hands-on experience with Prometheus, Alertmanager, and Grafana, including metric collection, alert rules, routing, and dashboards; experience with VictoriaMetrics would be a plus;
- Experience with Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards for log collection, transport, processing, storage, search, and visualization;
- Understanding of metrics and log pipeline design, including reliability, scalability, data retention, capacity planning, and cardinality management;
- Experience designing actionable alerts, reducing alert noise, and troubleshooting infrastructure and application issues using metrics and logs;
- Proficiency in shell command line usage, scripting, and automation tools like Ansible/Terraform;
- Python and bash scripting skills;
Apply
#remote