Job Description & Details
This is a solid full-stack role that leans heavily into the current generative AI wave, requiring you to bridge traditional web development with modern LLM integrations. If you are tired of building standard CRUD apps and want to actually ship production RAG pipelines and agentic workflows, this opening in Santa Clara gives you that exact sandbox.
What You'll Actually Be Doing
Your day-to-day will be a mix of spinning up responsive frontend interfaces in React and engineering robust backend APIs using Node.js or Python. Beyond standard web dev, you will spend a good chunk of your time figuring out how to reliably hook up LLMs like OpenAI or Anthropic, dealing with token limits, structuring prompt chains, and wiring up vector databases for semantic search. You will constantly balance keeping the UI snappy while managing asynchronous AI inference calls that do not always behave deterministically.
The Core Tech Stack
You need to be fluent in TypeScript/JavaScript for the frontend, alongside a solid backend language like Node.js or Python. But the real gatekeepers here are the AI tools: you need genuine hands-on experience with vector databases, Retrieval-Augmented Generation (RAG), and orchestration frameworks like LangChain or Semantic Kernel. Knowing your way around cloud platforms like AWS or Azure and containerizing your services with Docker is non-negotiable since these AI apps demand serious infrastructure.
Interview Expectations
Expect the hiring team to grill you on how you handle failure states when calling external LLM APIs, especially regarding latency and rate limiting. They will likely ask you to whiteboard a RAG architecture and defend your choice of vector database versus traditional SQL indexing. The interviewer is secretly looking to see if you understand the actual bottlenecks of generative AI in production, rather than someone who has only played with toy prompt scripts.
Application Advice
Skip the generic resume summary and put your AI integrations right at the top. Make sure to explicitly mention every time you've built a feature using OpenAI APIs, RAG, or vector search, as the ATS is heavily weighted toward these keywords. If you have deployed any agentic workflows or SaaS apps leveraging LLMs to production, highlight the scale and performance metrics to immediately catch the engineering manager's eye.