Job Description & Details
This is a heavy-hitting evaluation engineering gig focused entirely on agentic AI systems within regulated financial environments. You won't just be playing with prompts here; you will own the end-to-end evaluation infrastructure, building the safety nets and scoring architectures that make autonomous agents viable in production.
What You'll Actually Be Doing
Your day-to-day will revolve around designing offline regression suites, setting up online monitoring, and running A/B tests for complex multi-agent workflows. You'll spend a significant amount of time building LLM-as-judge and Agent-as-judge pipelines that can accurately assess trajectory quality, tool-call precision, and model safety. The real challenge here isn't just writing the Python code, but architecting these evaluation frameworks so they can handle enterprise-scale delivery—affecting dozens of applications while keeping strict compliance and audit requirements in mind.
The Core Tech Stack
You need to bring advanced Python and a solid 5-7 years of background building production-grade machine learning or generative AI systems. The foundational stack relies heavily on AWS Bedrock (including Models and Guardrails), CloudWatch, Lambda, S3, and IAM. On the AI side, deep hands-on experience with agentic frameworks like LangGraph, CrewAI, or similar is non-negotiable, along with a firm grasp of core evaluation metrics like task completion, groundedness, and trajectory quality.
Interview Expectations
Expect the hiring team to grill you on how you handle edge cases in LLM-as-judge pipelines, particularly around prompt drift and calibration drift over time. They will likely ask you to whiteboard a system architecture for continuous online monitoring of a multi-agent workflow that occasionally makes recursive tool calls. They aren't just looking for a textbook definition of evaluation metrics; they want to see how you balance thorough safety and compliance checks with low latency and real-world API rate limits.
Application Advice
To get past the ATS and catch the engineering manager's eye, make sure your resume explicitly highlights your experience with agent evaluation frameworks, offline regression testing, and specific metrics like tool-call accuracy and groundedness. If you have worked with AWS Bedrock or built custom LLM-as-judge scoring rubrics in past roles, bring those projects to the very top of your experience section. Don't bury your Python and cloud architecture skills under generic bullet points—quantify your impact where possible, especially if you have scaled evaluation pipelines in regulated domains like financial services.