Back to Jobs

Python Developer - AWS, AI, LLM

Not Disclosed

Job Description & Details

This role is a unique hybrid data engineering and GenAI position where you won't just be moving data around, but actually building the data pipelines that feed modern LLM and RAG architectures. If you like bridging the gap between traditional data engineering and bleeding-edge AI applications while sitting on-site in the DC area, this is worth a serious look.

What You'll Actually Be Doing

You'll spend your days designing and optimizing heavy-duty Python and SQL pipelines that prep messy enterprise data for consumption by large language models. Half the job is classic ETL/ELT work—building reliable data flows, mapping schemas, and dealing with AWS data stores—while the other half is deeply focused on GenAI enablement, crafting embeddings, hooking up vector databases, and ensuring your RAG setups don't hallucinate all over the production environment.

The Core Tech Stack

You need absolute fluency in Python and advanced SQL to survive here, along with hands-on AWS chops spanning Lambda, Redshift, and Glacier. Beyond the standard data stack, familiarity with Amazon Bedrock, vector databases, and prompt engineering is mandatory because you'll be actively integrating LLMs and building semantic search capabilities that business stakeholders actually depend on.

Interview Expectations

Expect them to grill you on how you handle data leakage and context window limitations when building RAG pipelines, likely asking you to whiteboard a flow from raw enterprise documents all the way to vector embedding generation and retrieval. They are also going to test your Python and SQL optimization skills under pressure, looking to see if you can debug an inefficient pipeline or explain how you measure hallucination rates and groundedness in production LLM outputs without sounding like you just read a blog post.

Application Advice

Make sure your resume doesn't just look like a standard software developer template; highlight the exact intersection of your data engineering background and any AI or LLM projects you have touched. Explicitly list out keywords like 'Retrieval-Augmented Generation', 'ETL/ELT', 'Python', 'AWS Lambda', and 'vector databases' so you breeze past the initial automated filters.