Job Description & Details
This is a heavy-duty, hands-on contractor role for someone who knows how to keep complex AWS environments and data pipelines running smoothly without hand-holding. You will be owning the infrastructure delivery, automation, and production reliability for their data platforms, meaning you need to be comfortable acting as the primary escalation point when things break. If you love building robust Terraform modules and untangling messy production incidents, this is a solid gig.
What You'll Actually Be Doing
Your day-to-day will be split between writing clean Infrastructure as Code, troubleshooting tricky production networking and container issues, and keeping an eye on the reliability of their data workflows. You'll spend a lot of time reviewing Terraform plans, managing IAM policies, optimizing AWS costs, and figuring out why a Redshift or Airflow pipeline failed at 2 AM. You won't just be designing systems on a whiteboard; you'll be deep in the trenches handling drift management, container operations, and disaster recovery testing.
The Core Tech Stack
AWS is the absolute foundation here, and they need someone with genuine production mileage in VPCs, ECS/Fargate, Lambda, S3, and RDS—surface-level exposure won't cut it. You also need rock-solid Terraform skills for writing reusable modules and handling state safely through CI/CD pipelines. On top of that, expect to lean heavily on your Linux networking fundamentals, Docker container operations, and Python or shell scripting for automating away manual operational friction. While you don't need to be a hardcore data engineer, familiarity with Redshift, Databricks, Airflow, and dbt is essential for keeping their data workloads healthy.
Interview Expectations
Expect them to grill you on real-world failure scenarios, particularly around Terraform state drift and troubleshooting a failing CI/CD deployment during a critical incident. The hiring manager is secretly looking to see how you balance moving fast with maintaining strict security practices like least-privilege IAM and secrets management. You'll likely also be asked how you approach diagnosing a networking bottleneck between a VPC and an RDS instance, so be ready to talk through your exact debugging methodology step-by-step.
Application Advice
To get past the ATS, make sure your resume explicitly highlights your hands-on experience with Terraform modules, AWS networking, and production SRE work. Don't just list technologies; write out bullet points that emphasize scale, uptime, and how you used Python or shell scripting to eliminate manual toil. If you have direct experience supporting Airflow, dbt, or Redshift infrastructure, bring those to the forefront to show you can handle the specific data platform requirements of this role.