Back to Jobs

QA Engineer for AI Products/Solutions

Not Disclosed

Job Description & Details

This is a heavy-hitting QA role specifically tailored for engineers who have moved past traditional web app testing and into the messy, probabilistic world of Large Language Models and Generative AI. With a 12-year experience requirement, you won't just be clicking buttons; you'll be defining quality standards for systems that constantly change their output. If you are tired of standard CRUD app regression testing and want to break models for a living, this Bellevue-based onsite position puts you right on the front lines of AI reliability.

What You'll Actually Be Doing

Expect to spend your days knee-deep in unpredictable model behaviors, prompt variations, and data pipelines. You'll be designing test plans that account for hallucinations, bias, and edge-case failures rather than just checking if a button returns a 200 OK. A massive part of your day-to-day involves building out automated test suites, running adversarial red-team exercises to test safety boundaries, and creating golden datasets to benchmark model performance across different version upgrades.

The Core Tech Stack

Python is non-negotiable here because you'll be writing custom scripts and automation frameworks to test complex backend services and APIs. You need deep familiarity with AI evaluation frameworks and SQL, as validating data pipelines, checking schema validity, and tracking data drift are central to keeping these models fed with clean data. Without a solid understanding of how LLMs consume prompts and spit out tokens, you'll struggle to define the right acceptance criteria alongside the data science team.

Interview Expectations

When you sit down with the engineering team, expect them to ask you how you would systematically test a model for hallucination and semantic drift without a deterministic expected output. They want to hear about your methodology for creating golden datasets and how you handle subjective metrics like 'accuracy' in a non-deterministic environment. Secondly, be prepared for a deep dive into your Python automation architecture and how you've scaled test suites for high-latency AI services under heavy load.

Application Advice

To get past the ATS and grab the hiring manager's attention, make sure your resume heavily features terms like LLM/GenAI product testing, AI evaluation frameworks, and adversarial red-teaming. Don't just list old Selenium or web-testing projects; highlight instances where you dealt with probabilistic outputs, data pipelines, and prompt engineering regressions. Show them you understand that testing an AI product is fundamentally different from testing traditional software.