In April 2025 a support AI at the developer tool Cursor told users about a multi-device login restriction that did not exist. It had not been trained to lie. Retrieval returned nothing authoritative and the model filled the gap with plausible-sounding language — which is what these models do when asked a question they cannot answer from source.
The failure is upstream of the model
It is tempting to treat an invented policy as a model defect, and to fix it by swapping models. That almost never works, because the model behaved normally. It was asked a direct question, had nothing authoritative in front of it, and produced the most probable-looking answer. A better model produces a more convincing wrong answer.
The actual failure happened one step earlier: the retrieval layer found nothing relevant and passed the question through anyway. Everything useful you can do about hallucination happens at that step, or in what the model is told to do when that step comes back empty.
Which is why the symptom clusters
Merchants report it in specific places, and they are always the places with no source to retrieve: users of knowledge-base bots report answers degrading after the first few queries and bots that ignore uploaded data or invent answers. Returns windows, delivery estimates, stock, whether a product suits a use case. These are the questions customers ask most and that most stores have never written down.
Three things that actually stop it
1. Tell the model that not knowing is an acceptable answer
Most prompts implicitly demand an answer. Unless you state the opposite, the model will treat "I don't have that" as a failure to be avoided. It has to be given explicit permission — and explicit instruction — to refuse.
The wording that matters is narrow and concrete rather than general. "Be accurate" does nothing. What works is naming the specific categories where invention is most likely and most costly: never state a price, stock level, availability, delivery time, policy, order status, size, measurement or material that is not in the source material or a tool result. Then say what to do instead, which is to say so plainly and offer to check with a person.
2. Give it live data for anything that changes
Pasting your catalogue into a knowledge base as text works for about a week. Then a price changes, something sells out, and the agent quotes last month's number with total confidence. Stale and fragmented context is the most common reason support agents fail in production — not the model.
Prices, stock and order status should be read at the moment of asking, from the store, through a tool. Anything read live cannot go stale between the sync and the customer.
3. Make the refusal lead somewhere
A bot that says "I don't have that information" and stops is only marginally better than one that invents. The single most common AI support failure is the absence of a clear, fast path to a human, and it is the failure underneath most of the others. The refusal has to hand over — with the conversation attached, to somebody who is actually there.
What this looks like in practice
Thredo's agents run with a set of hard limits that sit above anything a business writes in their own instructions. They may state only what appears in the business facts, the knowledge, or a tool result. They may not invent or estimate a price, stock level, availability, delivery time, policy or order status. They may not say an action succeeded unless the tool returned success. And text that arrives from a tool or a document is treated as data, never as instruction — so a scraped page cannot tell the agent what to do.
The practical effect is that the agent is often less impressive in a demo and considerably more useful on a Tuesday. It will tell a customer it does not know your returns window rather than guessing at 30 days, which is the right trade for a business that has to honour whatever it said.
The part you cannot outsource
None of the above fixes a knowledge base that does not contain the answer. The most common sequencing mistake is choosing a tool before writing the knowledge — tool selection takes a day, and the writing takes weeks and decides the outcome.
The good news is that the list is short and finite. Go through your last fifty conversations, find the twenty questions that account for most of them, and write down the real answers. That document is worth more than any model choice, and it is the one piece of work no vendor can do for you.
How to test a vendor on this in ten minutes
Ask their agent something your knowledge base definitely does not cover. A niche policy question works well: whether you ship to a country you do not, or what happens to a return after ninety days.
A good agent says it does not have that and offers a person. A bad one produces a confident, specific, wrong paragraph. It takes one question, and it tells you more than the feature list.
Common questions
Why do AI chatbots make up policies?
Because the retrieval step found nothing authoritative and the model was still asked to answer. Language models produce the most plausible continuation when they have no source, so a missing returns policy becomes an invented one. The failure is in retrieval and instruction, not usually in the model itself.
Does using a better model fix hallucination?
Rarely. A stronger model produces a more convincing wrong answer when it has nothing to work from. What fixes it is giving the agent live data for anything that changes, explicit permission to say it does not know, and a fast handover to a person.
How do I test whether a support bot hallucinates?
Ask it something your knowledge base definitely does not cover, such as shipping to a country you do not serve. A well-built agent will say it does not have that and offer a human. A poorly built one will produce a confident and specific wrong answer.
What should I write down before setting up an AI agent?
The twenty questions that account for most of your conversations, with the real answers. Delivery times, returns window, sizing, what you do and do not ship. Tool selection takes a day; this takes weeks and decides whether the agent works.
Published 13 September 2026. The Cursor incident is widely reported and dated April 2025; competitor behaviour described here is drawn from public reviews, linked at each mention.