Back to Jobs

Machine Learning Engineer (Python and SRE Focus)

Not Disclosed

Job Description & Details

This is a heavy-duty production reliability role tailored for someone who understands machine learning workflows but lives and breathes Site Reliability Engineering. If you enjoy wrangling on-premises servers, keeping Kubernetes clusters healthy, and making sure ML models don't crash in production, this is right up your alley. It is an onsite gig in Plano, TX, so you will need to be local and ready for an in-person interview loop.

What You'll Actually Be Doing

You'll spend your days keeping machine learning applications humming across a hybrid mix of Windows and Linux environments, primarily hosted on-premises rather than pure cloud. Your main job is bridging the gap between data science and ops—taking the models your team builds and making sure they deploy, scale, and monitor without breaking. You will heavily utilize DataDog to set up alerting, dig into production logs when things inevitably go sideways, and automate the operational grit so you aren't doing manual deployments all day long.

The Core Tech Stack

You need absolute fluency in Python, as you'll be writing scripts for automation, debugging, and maintaining ML apps. On the infrastructure side, deep hands-on experience with Kubernetes, Docker, and managing bare-metal or on-prem servers running both Linux and Windows is non-negotiable. Familiarity with DataDog for observability and a solid grasp of CI/CD pipelines specifically tailored for ML workloads will keep you from drowning in this role.

Interview Expectations

Expect the hiring team to grill you on your incident management approach, especially regarding distributed systems and Kubernetes pod failures. They will likely ask you to trace a scenario where a deployed ML model experiences a memory leak or severe latency spike in production and walk them through your debugging steps using Python and DataDog. The interviewer is secretly looking for your calm under pressure and whether you treat infrastructure with the same rigor as application code.

Application Advice

Make sure your resume leans hard into your SRE, DevOps, and infrastructure background rather than just focusing on model training. They explicitly do not want pure Data Scientists or MLOps theorists; they want someone who can configure an on-prem cluster and debug a stubborn Linux or Windows server issue. Highlight keywords like Kubernetes, DataDog, on-premises infrastructure, Python automation, and CI/CD pipelines right at the top to clear the automated filters.