Back to Jobs

HPC / Kubernetes Consultant

Not Disclosed

Job Description & Details

This is a heavy-hitting infrastructure gig bridging the gap between traditional High-Performance Computing and modern cloud-native systems. If you like untangling gnarly cluster networking issues and squeezing every drop of performance out of GPU nodes for AI workloads, this role will keep you on your toes.

What You'll Actually Be Doing

You will spend your days wrangling massive Kubernetes clusters—we are talking 500+ nodes—and making sure GPU-heavy AI and ML pipelines don't bottleneck. Your time will be split between automating infrastructure with Terraform, debugging elusive scheduler or CNI failures, and optimizing Slurm and AWS environments like FSx for Lustre. Expect to handle incidents under pressure and lead root cause analyses when production clusters decide to misbehave.

The Core Tech Stack

You absolutely cannot fake your way through the Kubernetes requirements here; you need deep, production-level scars from managing large clusters, troubleshooting core components, and handling upgrades without dropping traffic. Paired with that, you need solid AWS chops (EKS, EC2, VPC, IAM), hands-on experience with NVIDIA GPU scheduling, and fluency in Slurm for HPC workload management. Python and Terraform glue the whole operation together, so your automation game better be airtight.

Interview Expectations

Expect the hiring team to grill you on cluster resilience by asking how you would debug a scenario where GPU pods are failing to schedule due to a complex CNI or resource fragmentation issue in a 500-node cluster. They want to see your mental model for isolation, networking bottlenecks, and metric monitoring under load. Another favorite will likely involve tracing a massive performance drop in a distributed AI training job running over InfiniBand or high-speed storage, where they will test your ability to methodically isolate whether the bottleneck sits in the network, the GPU memory, or the scheduler.

Application Advice

To make it past the ATS, your resume needs to explicitly highlight numbers—mention cluster sizes, node counts, and scale, because vague references to 'managing Kubernetes' won't cut it here. Front-load your experience with NVIDIA GPUs, Slurm, and AWS infrastructure automation, and ensure keywords like Karpenter, Cilium, and FSx for Lustre are woven into your project descriptions if you have touched them.