Job Description & Details
This is a heavy-hitting SRE leadership role at Early Warning, the company keeping critical financial plumbing like Zelle and Paze running smoothly for hundreds of millions of users. You won't just be managing people from an ivory tower; they explicitly expect you to keep your hands dirty, write code, and contribute to automation when things get deep in the weeds. If you like high-stakes distributed systems and want to drive reliability at a massive scale, this is a solid place to plant your flag.
What You'll Actually Be Doing
You'll spend your days balancing high-level strategy with gritty production problem-solving across an assigned product domain. Your core mission is translating business priorities into measurable reliability outcomes by setting hard-hitting SLOs, managing error budgets, and ruthlessly hunting down toil. Expect to partner heavily with Product, Security, and Software Engineering to eliminate architectural bottlenecks, run blameless post-mortems that actually lead to change, and mentor a talented crew of engineers and managers.
The Core Tech Stack
While the job description keeps the exact tech stack a bit high-level, you need deep, battle-tested experience with cloud platforms, distributed architectures, and modern observability tooling at scale. Because this is a financial services ecosystem, mastery of infrastructure automation, CI/CD pipelines, and robust telemetry (metrics, logs, traces) is non-negotiable. You'll need to know how to build systems that fail gracefully and recover instantly under heavy transaction loads.
Interview Expectations
Expect the hiring team to probe heavily into your incident management philosophy and how you handle architectural tradeoffs between velocity and reliability. They'll likely ask you to describe a time a critical system failed under your watch, how you handled the post-mortem without slipping into a blame culture, and what concrete code or automation you introduced to prevent it from happening twice. They want to see that you can mentor junior leaders while still dropping into an IDE or terminal to inspect code or review an automation script yourself.
Application Advice
To get past the ATS, your resume needs to scream measurable impact rather than just listing job duties. Make sure to weave in keywords like "SLOs/SLIs", "error budget management", "distributed systems observability", and "infrastructure automation". Don't just say you managed SRE teams—quantify it by highlighting the percentage by which you reduced toil, scaled throughput, or improved system availability during high-traffic surges.