Job Description & Details
This is a heavy-hitting L6 architecture gig focused on building out enterprise-grade agentic infrastructure from scratch. You will not be writing simple wrapper scripts; instead, you are defining the foundational patterns for how autonomous agents execute, interact with models, and safely call enterprise tools at scale. If you are tired of building toy demos and want to solve the actual production nightmares of LLM observability, Zero Trust security, and state management, this is the table you want to sit at.
What You'll Actually Be Doing
Your day-to-day will be split between high-level systems design and hands-on standardization of agent execution paths. You will be architecting AgentCore gateways, mapping out memory tiers between vector stores and operational databases, and implementing strict human-in-the-loop controls. Expect to spend a lot of time arguing trade-offs between framework flexibility—like LangGraph, OpenAI Agents SDK, and Google ADK—and maintaining strict enterprise governance, blast-radius containment, and cost efficiency across millions of token calls.
The Core Tech Stack
AWS is non-negotiable here, and you are expected to bring active AWS certifications to the table. Beyond cloud infrastructure, you need deep, production-level familiarity with modern agent frameworks, Model Context Protocol (MCP), and Agent-to-Agent (A2A) standards. The team relies heavily on robust vector databases, distributed state management systems, and comprehensive AgentOps tooling for tracing, latency monitoring, and token cost tracking. If you haven't engineered secure identity propagation and Zero Trust boundaries for LLM tool-calling in the past, you will struggle with the security demands of this architecture.
Interview Expectations
Expect the interview panel to grill you on failure modes in distributed agent loops, specifically asking how you handle infinite execution loops, runaway token costs, and silent tool-calling failures. They want to see how you design blast-radius containment when an autonomous agent hallucinates a destructive API call. Another classic discussion point will be memory architecture—be ready to defend your strategy on when data should live in short-term context versus long-term vector stores or enterprise systems of record without violating privacy or token limits.
Application Advice
Skip the generic AI buzzwords and tailor your resume to highlight your actual architectural battle scars with LLMs and distributed systems. Explicitly mention any production experience you have with AWS, AgentOps, framework integrations like LangGraph or custom SDKs, and enterprise security frameworks like Zero Trust and MCP. Make sure your active AWS certifications are front and center at the top of your resume so you immediately clear the non-negotiable HR filter.