Systems Engineer (Core Infrastructure)
iForce Connect ·
- Work mode
- Remote
- Category
- Software engineering
iForce Connect ·
9X5 Consulting · Melbourne VIC
The GMP Group · Singapore
Dell · Mclean, VA, United States
NXP Semiconductors · Eindhoven
Remote systems engineer who designs, deploys, and scales a customer-facing cloud GPU compute platform: bare-metal provisioning, GPU cluster and CUDA management, KVM/Kubernetes virtualisation, low-latency networking (BGP, SDN, VPC), distributed storage, and automation with Terraform/Ansible to support AI/ML workloads.
This is a remote position.
We are seeking an experienced Systems Engineer to design, deploy, and scale our high-performance cloud compute platform and distributed GPU infrastructure. In this role, you will be responsible for building and maintaining bare-metal hardware, high-density GPU clusters, virtualisation layers, and low-latency networking that power complex AI/ML workloads. You will work closely with Software, Architecture, and Product teams to drive the performance, reliability, and security of our customer-facing compute infrastructure.
Provision, manage, and optimise large-scale bare-metal servers and specialised hardware configurations across multi-datacenter environments.
Architect and maintain high-performance compute clusters equipped with modern GPU accelerators, managing driver deployments, CUDA runtime environments, and firmware life-cycle operations.
Build and automate hypervisor environments (KVM, QEMU) and container orchestration platforms (Kubernetes) tailored for intensive parallel computing and multi-tenant isolation.
Implement and maintain ultra-low-latency network fabrics, Virtual Private Clouds (VPCs), and software-defined networking (SDN) solutions.
Deploy and scale high-throughput distributed storage architectures designed for data-heavy training and inference pipelines.
Optimise system performance across storage, memory, compute, and inter-node interconnects to ensure maximum platform utilisation.
Develop infrastructure-as-code (IaC) configurations using tools such as Terraform, Ansible, or custom automation scripts to ensure rapid, reproducible bare-metal provisioning.
Design proactive monitoring, alerting, and observability frameworks to track system telemetry, hardware health, and thermal efficiency.
Participate in high-availability architecture planning, disaster recovery testing, and incident response to maintain stringent SLA targets.
Evaluate emerging server hardware, accelerator technologies, and cloud orchestration frameworks to continuously refine system design.
Collaborate with Software and Platform Engineering teams to build seamless APIs and control planes on top of raw physical infrastructure.
Maintain comprehensive architecture documentation, hardware baseline specs, and operational runbooks for complex infrastructure workflows.
Linux Kernel & Systems Administration: Deep, hands-on mastery of Linux systems (Debian/Ubuntu, RHEL/Rocky, Arch variants), kernel tuning, and low-level system performance optimisation.
Hardware & Accelerators: Proven experience managing enterprise server hardware, high-density server chassis, and GPU/accelerator acceleration platforms.
Virtualisation & Containers: Strong proficiency with Kubernetes, Docker, KVM, and cloud management frameworks (e.g., OpenStack, custom orchestrators).
Automation & IaC: Expertise in automated system provisioning, Configuration Management (Ansible, Puppet, or Chef), and Infrastructure-as-Code (Terraform).
Networking & Security: In-depth knowledge of BGP, VLANs, overlay networks, firewall policies, zero-trust architectures, and multi-tenant security isolation.
Strong analytical and troubleshooting skills under high-pressure, live operational conditions.
Passion for high-performance computing, open infrastructure, and scalable system design.
Clear communication skills with a proven ability to collaborate across software, architecture, and operational disciplines.
Dell · Mclean, VA, United States