Back to Jobs

Site Reliability Engineer - (GPS-Salesforce CRM & YAVA)

Not Disclosed

Job Description & Details

This is a heavy-duty, hands-on SRE contract focused entirely on preventing production fires across critical Salesforce CRM and YAVA servicing platforms. You aren't just writing Terraform or answering pages here; your primary mandate is building the synthetic monitoring, dependency maps, and automated validation to catch degradation before it hits agents or customers. If you like owning the holistic health of complex distributed systems and squeezing out toil, this one's worth looking at.

What You'll Actually Be Doing

Your days will be split between writing proactive detection logic, building synthetic probes for agent and customer journeys, and tearing down postmortems when things go sideways. You'll spend a lot of time mapping out tricky dependencies—everything from DNS, TLS certificates, and API gateways down to legacy IBM MQ and mainframe/DB2 connections. When high-risk changes drop, you're the gatekeeper defining pre-and-post change validation, running load tests, and ensuring circuit breakers and timeouts are actually configured to prevent cascading failures instead of making them worse.

The Core Tech Stack

You need a deep foundation in SRE tooling and distributed systems, specifically observability stacks (metrics, logs, traces, synthetic monitoring), and robust scripting skills in Python, Go, or PowerShell to automate away operational toil. Because of the platform scope, hands-on exposure to Salesforce CRM, CTI integrations, and contact center technologies like telephony or conversational AI will make or break your ramp-up time. Networking fundamentals—DNS, load balancers, WAFs, API gateways, and TLS certificates—are non-negotiable since so much of this role revolves around shared dependency health.

Interview Expectations

Expect the hiring team to push hard on your incident response and debugging methodology, likely walking you through a hypothetical cascading failure involving a third-party vendor outage or a slow database connection pool. They want to see how you trace root causes across distributed components when logs are sparse and telemetry is noisy. You'll also need to defend your approach to synthetic monitoring and chaos engineering, so be ready to explain how you design safe production probes that don't accidentally trigger alerts or corrupt live data.

Application Advice

To get past the ATS, your resume needs to explicitly highlight measurable outcomes around incident reduction, MTTD/MTTR improvements, and automation wins rather than just listing tools. Mirror the technical vocabulary from the job description by spelling out your experience with SLIs/SLOs, change validation frameworks, and complex dependency troubleshooting. If you've worked with Salesforce, contact center tech, or heavy enterprise integrations like MQ and DB2, bring those to the absolute top of your experience section.