I was troubleshooting why our Email Triage Digest, a scheduled task that drafts email replies each morning, did not always run on schedule. I asked Claude for help. It found a reasonable cause and suggested a fix, then had a good idea of its own: add a notification for when the task fails, so we would know right away instead of noticing by absence. I asked how to turn that on. Claude told me exactly where to click.
I went there. The option was not there.
I described what I actually saw, and Claude was confident again: it must be a general setting somewhere else. I checked. Nothing. I asked where else to look, and it named three more places. I checked all three, screenshots and all. None of them had it either. Only when I stopped following directions and asked directly why it had sent me looking for something that did not exist did it stop guessing and actually verify. The setting does not exist yet, not anywhere in the product. Claude was not lying. It genuinely believed the functionality was there.
This is not a story about Claude being unreliable. It is a story about how close I came to just believing it anyway, the same way a team believes a confident technical leader who never actually checked either.
Confidence Is Not the Same as Correctness
An AI model is very good at generating a plausible next step. Plausible is not the same as verified. Stanford’s 2026 AI Index Report found that 74 percent of organizations now name inaccuracy as their top AI risk, up fourteen points in a single year, ahead of cybersecurity, regulatory compliance, and privacy combined. The same report found that even the best-performing models still return an incorrect answer roughly one time in five.
That is not a rare, unlucky day for a model. It is a known, measured, common experience across the industry, which is exactly why the response to it needs to be built into how a person works with the tool, not treated as a surprise. The risk is highest in exactly the moment I was in: a specific, technical, “click here” instruction that sounds like it came from someone who already checked. There was no hedge in the answer, no “I believe” or “this should work.” It read like a fact.
What Actually Went Wrong, In Claude’s Own Words
When I finally asked directly why it had pointed me to a setting that did not exist, Claude gave me a straight answer instead of another guess. It had not checked the actual configuration schema before answering the first time. It assumed a control existed because that is the kind of control most tools like this one usually have, and it told me where to find it instead of confirming first. When I said nothing was there, it guessed again, three more times, before it finally went and checked through the tool itself rather than trusting its own assumption.
That is the mechanism worth remembering. An AI model answering a “how do I” question is often completing a pattern from what a similar system usually looks like, not reading the actual system in front of it, unless something in the conversation forces it to check first. I have watched leaders make the exact same move with people instead of systems.
Technical Leaders Do the Same Thing
It happens constantly. A former engineer now running a program, someone with enough real depth to sound completely credible, tells the team “that capability is already there” or “just point them at the dashboard, it’s built in.” They are usually right, which is exactly what makes the moments they are wrong so expensive. Nobody downstream questions a confident instruction from someone who clearly knows the domain, the same way I did not question Claude’s confident click-here instruction.
A leader who assumes a system already does something, then directs the team to build on that assumption, can send a whole sprint down the same dead end Claude sent me down for fifteen minutes. The team does not push back. They assume they are missing something, the same way I first assumed I was clicking in the wrong place, and burn a day looking for functionality that was never there. The fix is not less confidence, which is usually earned and genuinely useful. The fix is treating a specific factual claim, “this exists,” “this is already configured,” as something to confirm before it becomes an instruction, not after someone has already spent a day acting on it.
The Habit That Actually Caught It
What caught this was not politeness. It was screenshots. It was asking exactly where the setting supposedly lived, more than once. And it was eventually asking directly: why did you send me somewhere that does not exist? That question is what produced a real answer instead of a fourth guess.
This is exactly the discipline our AI Training program for technical program managers and Scrum Masters is built around. AI can own drafting and even diagnosing a problem, but confirming a specific factual claim against the real system is a judgment call, not something to hand off. The program’s closing session teaches output verification as a formal habit, not a personality trait, because leaving it to individual instinct means some people catch it and some do not. An organization figuring out where that habit is missing, in how its people use AI or how its leaders direct their teams, can start with our AI Readiness Self-Assessment, which scores this kind of process maturity before it becomes an expensive surprise.
The frustrating part was never that Claude got something wrong. It was how easy it would have been to stop one step earlier: take the first confident answer, try it once, and assume I was doing something wrong instead of questioning the answer itself. Most people troubleshooting on a busy morning would have tried two or three of the suggested locations, not found the setting, and quietly concluded they were bad at this, rather than that the instruction was wrong. The actual problem, the digest failing silently, got fixed and confirmed working in that same conversation. The side trip chasing a setting that does not exist is the part worth remembering.
I still use Claude the same way I did before that morning, and I still trust the technical leaders I work with. What changed is that a specific, confident instruction, whether it comes from a model or from someone above the actual work, now gets one extra step before I act on it or pass it down. I check it against what is actually there, not what someone assumed was there.
If your teams are using AI daily, or your leaders are directing work based on assumptions nobody has checked lately, that gap is exactly what we help close in the AI adoption work we do at Stephens Insight Group. Talk to us about your team’s AI adoption.
Sources
- The 2026 AI Index Report (Stanford HAI)