Back to Jobs

AI Engineer – RAG / LLM / Agentic AI

Not Disclosed

Job Description & Details

This is a specialized senior engineering role focused on building advanced generative AI systems using RAG and autonomous agents. If you have spent the last decade deep in software engineering and want to pivot into designing complex, production-grade LLM architectures, this is a solid opportunity to build cutting-edge systems rather than just wrapping APIs.

What You'll Actually Be Doing

You will spend your days architecting and optimizing retrieval-augmented generation pipelines and agentic workflows that can handle high-throughput, domain-specific tasks without hallucinating. Your real challenge won't just be calling LLM endpoints, but solving the messy engineering problems around context window limitations, latency, vector database scaling, and ensuring agents make reliable decisions in production environments.

The Core Tech Stack

You need absolute fluency in building RAG systems and working directly with Large Language Models, coupled with deep expertise in Agentic AI frameworks. Because this is a high-level W2 role, the team expects you to understand how to bridge traditional enterprise backend systems with modern stochastic AI architectures, utilizing vector databases, embedding models, and multi-agent orchestration patterns effectively.

Interview Expectations

Expect the hiring managers to skip the softball questions and throw a complex architecture scenario at you, such as asking you to design a fault-tolerant multi-agent system that queries multiple unstructured enterprise databases via RAG. They are secretly testing whether you understand the actual failure modes of LLMs in production or if you just know the marketing buzzwords. They will also drill down into how you manage context retrieval latency and handle token budget optimization under heavy load.

Application Advice

Do not just list RAG and LLMs as bullet points on your resume; your application needs to prove you have spent 10+ years building and scaling complex systems, with a clear subsection highlighting production deployments of generative AI. Make sure to explicitly include terms like vector databases, context window optimization, multi-agent orchestration, and LLM fine-tuning or prompt engineering to easily clear the ATS filters.