Back to Jobs

ML Engineer - LLM Platforms & Assistants

Not Disclosed

Job Description & Details

This is a hands-on engineering gig focused on turning experimental LLM setups into production-grade systems on AWS. If you are tired of just writing prompt wrappers and actually want to build scalable RAG pipelines, manage infrastructure, and deal with enterprise-grade data security, this role is worth looking into.

What You'll Actually Be Doing

Your day-to-day will revolve around architecting and maintaining production LLM services that connect OpenAI models and emerging orchestration tools to enterprise data sources. You will spend a lot of time writing modular Python services, setting up Retrieval-Augmented Generation (RAG) workflows, and deploying everything securely onto AWS compute platforms like ECS, EKS, or Lambda while ensuring proper monitoring through CloudWatch.

The Core Tech Stack

Python is your bread and butter here, alongside deep hands-on experience with OpenAI integrations and LangChain or agentic workflows. You absolutely must know your way around AWS—specifically IAM, KMS, Secrets Manager, and storage services like S3—because shipping secure code is non-negotiable for this platform. Bonus points if you've already tinkered with Amazon Bedrock and AWS data tools like OpenSearch or Athena for semantic search.

Interview Expectations

Expect the hiring team to grill you on how you handle failure modes and latency bottlenecks in RAG architectures. They will likely ask you to explain how you would design a fault-tolerant caching layer for LLM responses to reduce API costs and improve throughput. The interviewer isn't just looking for a textbook definition; they want to hear about your battle scars regarding token limits, context windows, and keeping API keys secure in an enterprise environment.

Application Advice

Make sure your resume highlights your production-level Python and AWS deployments rather than just side projects or toy scripts. Explicitly call out any metrics where you successfully reduced latency, optimized vector database queries in OpenSearch, or managed AWS IAM roles securely. Weave keywords like RAG, semantic search, ECS/EKS, and LangChain naturally into your past work experience to get past the automated screeners.