When AI Performance Starts to Drift

What happens when the AI model you trust starts getting worse?

A conversation in the Promptmates community caught my attention this week.

Several TA leaders are noticing changes in how Claude is performing in real recruiting workflows.

One person said Claude, which had been their trusted model for assessing candidate fit, has recently started making outright mistakes and struggling to self-correct even after being flagged.

Another said Claude has started losing capabilities between runs. It could complete steps 1–5 one day, then suddenly claim it couldn’t complete step 1 the next. After being shown its own logs, it could suddenly do the task again.

Someone else described Opus as “lazy,” saying it was making assumptions, not checking its work thoroughly and running into errors on scheduled tasks.

And one person said they tested alternatives and found Gemini was currently producing more accurate and reliable candidate-fit assessments.

Then there was an interesting theory about why this might be happening.

One community member had just spent an entire day in training with Anthropic and noticed that Anthropic was pushing Opus 5.0 hard. His theory was that perhaps the newer model is more computationally efficient for Anthropic to run, even if customers are paying roughly what they paid for older models.

But the bigger takeaway is worth paying attention to: The newest model isn’t automatically the best model for your workflow.

Some people in the conversation are deliberately staying on older versions rather than moving to the newest release.

Others are finding Sonnet works better for their day-to-day work and only use a more advanced model when a task actually requires deeper reasoning. I personally use Sonnet for almost everything.

And there are people finding Gemini performs better for specific recruiting use cases.

There were also some practical ideas for correcting performance issues:

Audit your instructions. One person has been using Claude’s /doctor command to review everything they’ve built and identify overfitted instructions, unclear requirements, redundancies and bloated documentation.

Try giving the model less instruction. Anthropic has been pushing the idea of “instruct less and let Claude cook,” testing whether the model can accomplish more with fewer constraints and then adding guardrails where they’re actually needed.

And test models against your actual work.

If you’re using AI to score candidates, don’t decide which model is “best” based on benchmarks or what everyone on LinkedIn is talking about.

Take 20 candidates you’ve already evaluated manually.

Run those same candidates through Claude, Gemini and whatever other model you’re considering.

Compare the results.

Look at false positives. Look at false negatives. Look at consistency. Look at whether the model is making assumptions that a recruiter wouldn’t make.

Now you have something much more useful than an opinion about which model is better.

You have evidence.

AI models change. Their behavior can change. Your prompts can change. Your workflows can change.

So the model that was perfect for your recruiting workflow six months ago might not be the model you should be using today.

The question really is: “What’s the best model for this specific job, and how do we know?”

Never Miss a QA Post

Get the latest posts and tips delivered straight to your inbox.

I don’t spam! Read my privacy policy for more info.

Leave a Reply

Your email address will not be published. Required fields are marked *