A hands-on data center engineer (or lead) keeping high-density GPU server clusters running for an AI infrastructure and cloud services company in Southeast Asia: racking, cabling, hardware replacement, firmware, diagnostics, and 24/7 monitoring, with seniors also guiding junior technicians. Core stack includes Dell/HPE/Lenovo/Supermicro enterprise servers, Linux/Windows Server, and networking fund
Keep critical infrastructure running. Support the next wave of AI.
Join an AI infrastructure and cloud services company supporting large-scale GPU and data center deployments across Southeast Asia. We’re looking for a hands‑on engineer to keep high-density server clusters reliable, resolve technical issues, and support smooth deployments. For candidates with leadership experience, the role also includes guiding junior technicians and coordinating site activities.
What You’ll Do
Deploy and maintain servers: Handle rack‑and‑stack, cabling, hardware replacement, firmware updates, diagnostics, and preventive maintenance.
Troubleshoot critical issues: Diagnose server, storage, network, operating system, and connectivity faults; coordinate escalations with engineering teams and vendors.
Keep operations running: Monitor infrastructure health, alerts, capacity, and performance in a 24/7 environment.
Guide the team: Support junior technicians during deployments, maintenance windows, and incident response.
Own operational documentation: Keep asset inventories, spare parts records, incident tickets, SOPs, MOPs, EOPs, and shift handovers accurate.
Coordinate across teams: Work with data center operators, OEMs, network teams, logistics providers, and customers to resolve issues and deliver deployments.
Drive improvements: Streamline processes, automate routine tasks, and participate in shift coverage, on‑call support, and disaster recovery drills.
What You’ll Bring
3 to 10 Years experience in data center operations, server infrastructure, hardware maintenance, or field engineering.
Hands‑on experience with Dell, HPE, Lenovo, Supermicro, Inspur, H3C, or equivalent enterprise servers.
Strong practical knowledge of server hardware, storage, cabling, switches, optics, power distribution, and rack‑level troubleshooting.
Working knowledge of Linux and/or Windows Server, networking fundamentals, and diagnostic tools.
Experience with monitoring, ticketing, and asset management systems.
Clear communication, solid documentation skills, and a calm approach to resolving incidents.
A diploma or degree in IT, Computer Engineering, Electrical/Electronic Engineering, Networking, or a related field.
Stand Out With
Experience with NVIDIA GPU platforms, AI/HPC clusters, or high-density computing.
Exposure to large-scale deployments, commissioning, and automated OS installation.
Experience guiding technicians or acting as a site lead.
Familiarity with Cisco, Juniper, Arista, DCIM, or BMS.
Relevant certifications such as CCNA, Linux, ITIL, or CDCP.