Job Description & Details
This is a heavy-duty infrastructure gig for someone who lives and breathes massive Kubernetes clusters and GPU acceleration. If you are tired of standard web app deployments and want to tackle scaling 500+ node clusters, wrangling Slurm, and keeping AI/ML workloads humming without hitting bottlenecks, this role is worth a serious look.
What You'll Actually Be Doing
You will spend your days architecting, troubleshooting, and scaling complex environments where raw compute power is everything. Expect to debug deep scheduler issues, optimize low-latency networking with Calico or Cilium, and make sure that massive NVIDIA GPU training runs don't stall due to storage or node lifecycle issues. You are the backbone supporting data scientists and ML engineers, ensuring their distributed training pipelines have the infrastructure muscle they need.
The Core Tech Stack
You need absolute mastery over large-scale Kubernetes production environments, Slurm workloads, and NVIDIA GPU infrastructure. AWS is your playground here—specifically EKS, EC2, and FSx for Lustre for high-speed parallel file systems. Terraform and Python are non-negotiable for automating infrastructure out of trouble, and you'll need solid SRE chops with Prometheus and Grafana to catch latency spikes and cluster degradation before the pager goes off.
Interview Expectations
Expect them to grill you on cluster troubleshooting under load, such as how you would diagnose a cascading node failure in a 500-node EKS environment running heavy GPU workloads where pods are stuck in Pending due to CNI or storage driver timeouts. The hiring manager is secretly looking for your systematic debugging methodology and how you balance immediate remediation with long-term root cause analysis. They will also likely push you on InfiniBand, RDMA, or Slurm-to-Kubernetes migration strategies to see if you truly understand high-performance computing at the metal level.
Application Advice
Your resume needs to scream scale and performance right out of the gate. Do not just list Kubernetes and AWS; explicitly mention cluster sizes (e.g., managed 500+ node clusters), specific GPU models or scheduling setups, and concrete metrics around uptime or latency reductions. Weave in keywords like Slurm, FSx for Lustre, Terraform, and Cilium naturally into your past architectural achievements so the ATS flags you immediately as a senior infrastructure specialist who knows how to handle raw computing power.