Job Description & Details
This is a heavy-hitting leadership role for someone who knows how to move past the AI hype and actually build resilient, enterprise-grade execution planes. You'll be steering the infrastructure that powers everything from model gateways to complex agentic workflows in a high-stakes hybrid environment.
What You'll Actually Be Doing
Expect to spend your days bridging the gap between high-level architectural strategy and deep technical execution. You will lead teams building out model gateways, runtime environments, and retrieval pipelines while constantly fighting latency, scaling bottlenecks, and security compliance. A huge part of your time will go toward aligning disparate engineering groups, reviewing system designs, and making sure the platform doesn't buckle under the weight of enterprise-wide AI adoption.
The Core Tech Stack
You need a deep mastery of cloud-native infrastructure, specifically AWS ecosystems like EKS, Lambda, SageMaker, and Bedrock. Beyond the cloud primitives, absolute fluency in distributed systems, modern API design, and MLOps/LLMOps tooling is non-negotiable. They need someone who understands the nuances of vector databases, Retrieval-Augmented Generation (RAG) architectures, and multi-tenant isolation with robust RBAC.
Interview Expectations
They are going to grill you on how you handle failure states in distributed AI workloads, so expect a deep-dive scenario on designing a low-latency model gateway that can gracefully degrade under massive traffic spikes. The hiring manager is secretly testing your pragmatic trade-offs between strict enterprise security controls and developer velocity, making sure you don't build a fortress that nobody can actually ship code into.
Application Advice
Your resume needs to loudly showcase your experience leading engineering teams through major infrastructure overhauls, particularly those involving high-availability platforms and AI runtime environments. Make sure to weave in keywords like model gateways, distributed systems, CI/CD automation, and cloud-native observability to clear their ATS filters and prove you have the scars to lead this initiative.