Job Description & Details
This is a hands-on technical leadership role where you'll actually push production-grade GenAI systems instead of just building endless proof-of-concepts. You'll be bridging the gap between research and scalable software engineering, owning everything from vector databases to API orchestration. If you enjoy untangling ambiguous requirements and building resilient MLOps pipelines from scratch, this one's worth your attention.
What You'll Actually Be Doing
You'll spend your days designing and scaling end-to-end AI/ML pipelines, figuring out how to handle everything from data collection and preprocessing to real-time model validation. A major chunk of your time will go into building RAG architectures, handling long-context management, and wiring up agentic orchestration without blowing up your latency budgets. You'll also be setting up strict evaluation frameworks to monitor cost, drift, accuracy, and bias before pushing code to production via robust CI/CD workflows.
The Core Tech Stack
You'll need deep, production-level expertise in Python paired with heavy framework experience in PyTorch, TensorFlow, or Hugging Face. On the Generative AI side, expect to use LangChain or LlamaIndex alongside vector databases like Pinecone, Milvus, or Qdrant daily. Cloud infrastructure isn't optional here; you need to comfortably deploy and orchestrate workloads using Docker, Kubernetes, and managed services on AWS SageMaker, GCP Vertex AI, or Azure.
Interview Expectations
Expect to walk through a time you debugged a live production model suffering from severe data drift or runaway latency, where the hiring manager will look for your pragmatic trade-offs between cost, accuracy, and infrastructure overhead. They'll also likely hit you with a system design scenario around scaling a high-throughput RAG pipeline, checking if you truly understand context windows, token limits, and vector retrieval bottlenecks rather than just reciting textbook definitions.
Application Advice
Make sure your resume doesn't just list every tool you've ever touched, but instead highlights measurable metrics—like how you reduced model latency by 40% or scaled a vector search engine to millions of records. Explicitly sprinkle keywords like RAG, vector embeddings, MLOps, Kubernetes, and PyTorch throughout your experience bullets to effortlessly clear their initial ATS filters.