I grabbed Anantha, a QE leader for a conversation about AI, automation and technical hiring.
I went into the conversation with a pretty simple question: How the heck are recruiters supposed to vet technical talent when we don’t always understand what the engineering team is actually building?
I’ve been running into this more and more. Recruiters have asked me to help vet engineers, but when I ask about the actual engineering initiative, sometimes the answer is basically, “I need an automation engineeer for an AI adoption initative” or “Well Jaclyn, the team is moving to Playwright, so I need a Playwright SME.”
Okay… but why?
Let’s take Playwright as an example: If a team is moving from Selenium or Cypress to Playwright, my instinct as a recruiter might be to immediately go find someone with strong Playwright experience. But that’s not really enough information to understand the role.
Why are they migrating in the first place? What are they migrating from? Where are they at in the migration? What language are they using? Is the existing solution open source or commercial? Are we dealing with maintenance, test flakiness, licensing or something else?
Those answers change what the team actually needs from the person they hire.
And that was one of my first takeaways from the conversation: the technology isn’t necessarily the requirement. The business problem you’re trying to solve is.
Then we got into AI testing, and this idea became even more apparent.
Traditional application testing is relatively easier to wrap your head around. You have a requirement, you have an expected outcome, and you test whether the application did what it was supposed to do.
AI doesn’t work that way. The system may not produce the exact same response every time. Users can interact with it in ways you didn’t anticipate. And there may not be one predetermined answer you can simply compare against.
So how do you decide whether the response was actually good? And how do you build enough confidence to take something from a demo into production?
It’s not just about finding someone who has tested an AI product or used a particular AI testing tool. The bigger question is whether they understand how to evaluate something when the answer isn’t always binary.
That’s a much bigger skill than knowing one tool.
One concept that came up in our conversation was the idea of golden test sets. My understanding is that instead of assuming you can anticipate every possible scenario upfront, real-world interactions can surface new scenarios that can be reviewed, turned into test cases, and added to a growing set of examples you continue evaluating against.
That makes a lot of sense when you think about how unpredictable users can be. You can test the scenarios you expect people to encounter, and then someone inevitably does something you never thought of. Maybe they phrase a question differently. Maybe they give the system unexpected information. Maybe an agent takes an action you weren’t anticipating.
Now you’ve learned something and that scenario can become part of what you evaluate going forward.
So the testing process isn’t necessarily, “Build it, test everything, and you’re done.” There’s an ongoing feedback loop between what you expect the system to do and what actually happens when people use it.
And that brings me right back to recruiting.
If I’m looking for someone who understands AI evaluation, I don’t think the best question is simply, “What AI testing tools have you used?”
One of the most useful pieces of advice from the conversation was essentially to go one layer deeper than the resume.
For example, a candidate might list something like MCP on their profile. Great. But what did they actually do with it? How did they test it? What tool did they use? What were they looking for? Can they walk me through one specific test? What happened when something didn’t work? Get curious about the person and what they’ve done.
Those type of follow-up questions are much harder to answer with a polished, generic response.
And that’s really the point.
I’m not suggesting us recruiters turn into engineers, but I’m trying to get recruiters to stop treating technical keywords as proof of technical experience. There’s a difference between having worked with something and understanding what you were actually doing with it.
And AI makes that distinction even more important because candidates now have access to AI tools that can help them research terminology, generate answers and talk their way through concepts they may not have actually worked with.
So if I’m recruiting for an AI evaluation role, I want to get past, “Have you done AI testing?” and into, “Show me how you think about testing it.”
I came away from this conversation with a better appreciation for how much context recruiters can be missing when we start with the job description.
Selenium. Playwright. AI evaluation. MCP. Etc Etc.
Those words tell me something. They just don’t tell me enough.
The better questions are: Why are you doing this? What problem are you trying to solve? What does success look like? What does this person actually need to be able to do?
As recruiters, it’s important to understand the business initiative and day-to-day work well enough to ask better questions, understand what we’re recruiting for, and know when a candidate’s experience goes beyond the buzzwords.
The tool is rarely the whole story.
Huge nod to Anantha for taking the time to walk me through how he’s thinking about AI evaluation, automation and technical vetting. I learned a lot from the conversation and appreciate you letting me share some of those takeaways with the QE community. I highly recommend you check out his article on AI agents in production: https://www.ananthasubramanya.com/blog/enterprise-ai-evals
