Back to Jobs

AI Test Engineer

Not Disclosed

Job Description & Details

This is a specialized quality engineering role focused entirely on validating generative AI, LLM, and RAG systems within a regulated financial services environment. You won't just be writing standard UI regression scripts; you'll be building foundational evaluation pipelines to ensure AI model behavior, accuracy, and reliability at scale. If you want to bridge the gap between traditional QA and cutting-edge applied AI, this role gives you the playground to do it.

What You'll Actually Be Doing

You'll spend your days designing and deploying automated testing frameworks specifically built for LLMs and AI-enabled developer tools. Your core focus will revolve around creating evaluation pipelines that integrate directly into CI/CD workflows, allowing the team to continuously test for hallucinations, drift, and performance degradation. You will analyze complex, non-deterministic model outputs, trace token behaviors, and investigate why a RAG pipeline pulled irrelevant context, turning those findings into actionable metrics for stakeholders.

The Core Tech Stack

Python is non-negotiable here, as you'll be using it heavily for scripting custom evaluation logic and integrating with modern AI frameworks. You need a solid grasp of CI/CD environments and API testing, but the real differentiator is familiarity with AI-specific evaluation tooling like DeepEval, Ragas, or LangSmith. Experience working with cloud platforms like Azure AI, vector databases, and orchestration libraries like LangChain will dictate how quickly you can ramp up on building these automated test harnesses.

Interview Expectations

Expect the interviewers to probe your understanding of non-deterministic testing—specifically, how you mathematically or systematically validate an LLM's output when there is no simple 'pass/fail' string match. They will likely ask you to whiteboard a continuous evaluation pipeline for a RAG architecture, evaluating how you handle latency, token costs, and faithfulness metrics. They want to see that you understand the nuances of model risk management and can apply rigorous software engineering principles to chaotic AI outputs.

Application Advice

Your resume needs to heavily emphasize your Python scripting capabilities alongside concrete examples of building test frameworks from scratch. Make sure to sprinkle in keywords from the job description such as CI/CD integration, RAG systems, automated test suites, and evaluation pipelines to ensure you clear the automated ATS filters. If you have any open-source contributions or personal projects involving LLM observability or tools like LangSmith and DeepEval, feature them prominently at the top.