Back to Jobs

AI Engineer

Not Disclosed

Job Description & Details

This is a hands-on AI engineering gig focused on pushing Generative AI applications straight into production for entertainment platforms like DirecTV. You won't just be playing with notebooks here; you'll be building production-grade RAG pipelines and integrating LLMs into client applications across iOS, Android, and Smart TVs.

What You'll Actually Be Doing

Expect your day-to-day to revolve around building resilient, production-ready GenAI architectures using LangChain, LlamaIndex, and Python microservices. You'll be knee-deep in setting up complex RAG pipelines, fine-tuning open-source models like Llama and Mistral via LoRA, and ensuring user interactions are safe using guardrails like NeMo. Collaboration is huge here—you'll be working closely with MLOps folks to keep CI/CD pipelines flowing and monitoring token usage and inference latency in production.

The Core Tech Stack

Python is non-negotiable, and you need to be comfortable slinging code with FastAPI or Flask to build those crucial API endpoints. They rely heavily on LangChain and LlamaIndex alongside vector databases like Pinecone, Chroma, or Qdrant. Don't sleep on your containerization and cloud skills either—Docker, Kubernetes, and AWS or GCP experience are absolute requirements to keep these microservices humming.

Interview Expectations

Be ready to break down how you handle hallucination reduction in RAG pipelines; the hiring manager wants to hear about concrete strategies you've used to evaluate relevance and latency using frameworks like Ragas or TruLens. They'll also likely grill you on fine-tuning tradeoffs, specifically asking you to explain when you would choose LoRA versus full fine-tuning for an open-source model like Mistral.

Application Advice

Your resume needs to scream production-readiness, not just experimentation. Make sure keywords like LangChain, FastAPI, Pinecone, and QLoRA are front and center, backed by metrics showing how you scaled or optimized real-world AI applications rather than just academic projects.