Job Description & Details
The Data Sustain Lead role is a hands‑on position keeping a customer‑marketing data platform running day‑to‑day. You’ll own L2/L3 support, monitor pipelines, and act as the bridge between the engineering team and business stakeholders when incidents pop up.
What You'll Actually Be Doing
You’ll spend most of your time staring at pipeline health dashboards, digging into failed ETL jobs, and writing runbooks for recurring issues. Expect to triage alerts from Airflow and Snowflake, coordinate incident response in ServiceNow, and communicate RCA findings to on‑shore business owners. The job is as much about building reliable monitoring (CloudWatch, Grafana) as it is about writing quick Python or PowerShell scripts to patch data quality gaps before they affect downstream analytics.
The Core Tech Stack
The stack is a mix of classic ETL tools (SAP BODS, Informatica, IICS) and modern cloud components (Airflow DAGs, Snowflake, AWS Lambda/CloudWatch). You must be fluent in writing performant SQL on MS SQL Server for data reconciliation and comfortable scripting in Python, Bash or PowerShell to automate fixes. ServiceNow is the incident hub, so you’ll need to log changes, problems, and knowledge articles there. Grafana is used for visual monitoring, and a sprinkle of GitHub Copilot/Prompt Engineering shows they expect you to be comfortable with AI‑assisted code.
Interview Expectations
- “Walk me through how you would debug a failing Snowflake ETL job that’s also raising an Airflow DAG error.” – They want to see your systematic approach: checking Snowflake task history, reviewing Airflow logs, pinpointing where data quality checks failed, and how you’d use SQL and Python to isolate the bad records.
- “Explain a time you turned a noisy ServiceNow incident stream into actionable alerts. Which metrics did you track and how did you reduce MTTR?” – This probes your ability to design monitoring thresholds, write CloudWatch alarms, and document SOPs that cut resolution time.
Application Advice
Tailor your resume to surface the exact buzzwords: “BAU Data Pipeline Monitoring”, “L2/L3 sustain”, “ServiceNow Incident Management”, “SAP BODS”, “Informatica/IICS”, “Airflow”, “Snowflake”, “AWS Lambda”, “CloudWatch”, “Python/Shell scripting”, and “Grafana”. Highlight any runbook or SOP creation you’ve done and quantify impact (e.g., reduced MTTR by X%). If you have any Snowflake admin or dbt experience, put it front‑and‑center – it’s a nice‑to‑have that can push you ahead of the pack.