Back to Jobs

Senior Platform Engineer – Google Cloud Platform (DevOps & Agentic AI Infrastructure)

Not Disclosed

Job Description & Details

This is a heavy-hitting platform role where you'll be building the actual infrastructure to run production-grade multi-agent AI systems on GCP. If you enjoy wrestling with serverless cold starts, streaming web sockets, and zero-trust IAM for AI workflows rather than just managing standard CRUD app infrastructure, this one is worth a serious look.

What You'll Actually Be Doing

You'll be owning the entire request path from the browser all the way down to the Python-based AI agents and vector databases. Your day-to-day will involve balancing latency, cost, and security while setting up load balancers, configuring Cloud Run, and writing Terraform. You will work closely with AI engineers to ensure their multi-agent systems don't fall over in production, handling everything from zero-downtime CI/CD pipelines to distributed tracing for agent reasoning loops.

The Core Tech Stack

You need absolute fluency in Google Cloud Platform, specifically Cloud Run, Global HTTP(S) Load Balancers, and advanced IAM with Workload Identity. Terraform or OpenTofu is non-negotiable for managing multi-environment infrastructure as code. Since these are agentic AI workloads, experience with Python containerization, vector databases, and streaming protocols like Server-Sent Events or WebSockets will make or break your ability to keep these systems responsive.

Interview Expectations

Expect them to probe deeply into how you mitigate cold starts on Cloud Run while keeping infrastructure costs from spiraling out of control on high-throughput AI workloads. The hiring manager is secretly looking for your pragmatic trade-offs—can you balance absolute zero-trust security with the latency demands of real-time LLM streaming? You'll also likely be grilled on how you architect service-to-service authentication using ID tokens across complex agent graphs.

Application Advice

Make sure your resume doesn't just list generic DevOps keywords; put GCP front and center with exact metrics on scale, cost reduction, or latency improvements you've driven. Explicitly highlight your experience with Terraform modules, Cloud Run optimization, and any exposure you have to LLMs, vector databases, or AI agent runtimes to easily clear the ATS filters.