Stop AI making things up before it reaches your customers
Made-up answers are the biggest open problem in AI today. A model that does not know an answer rarely stays quiet. It offers something that sounds right, in the same confident voice it uses when it is correct.
That is the main reason many owners still will not put AI in front of their customers, and the caution is justified. This guide explains why it happens, how a system gets restricted to your content, and the five tests to run before you go live.
In this guide
Why does AI make things up?
AI makes things up because it is built to continue a sentence, not to verify a fact. When a detail is missing, it does not stop. It fills the gap with whatever sounds most like a correct answer.
Picture a quiz show contestant who loses nothing for a wrong answer and gains a point for a right one. That contestant will never say "I don't know." They will always say something, in a confident voice, because silence pays them nothing.
Why this is worse than an ordinary mistake
The problem is not only that AI gets things wrong, since people on a front desk get things wrong too. The problem is that a wrong answer looks identical to a correct one, so nobody checks it.
Your customer cannot tell the difference. If the system says a discount runs to the end of the month, they take that as your word. The damage is commercial, not technical.
Where guessing costs the most
Not every invented detail is equally expensive. It helps to know where the risk concentrates.
- Prices and discounts. A customer quoted a lower price expects you to honour it.
- Deadlines and cancellation terms. A wrong deadline creates a dispute nobody wins.
- Preparation instructions. In healthcare, a wrong instruction means a missed or unusable test.
- Availability. A promised slot that does not exist is worse than no answer at all.
Notice that all four are things no employee recites from memory either. A person would check. Unrestricted AI will not.
What a grounded answer means
A grounded answer is one the system may give only if it finds it in your content. No guessing from general knowledge, no blending your facts with something the model read elsewhere.
Back to the quiz. A grounded system is a student sitting an open-book exam. It may answer only what appears in the book you handed it. If the book is silent, it leaves the question blank.
The "I don't know" rule
This is the most important setting and the easiest one to test. A properly configured system has defined behaviour for the moment it has no answer.
- It admits the gap. No approximation, no improvising.
- It offers a next step. A contact route, a form, or a handoff to your team.
- It logs the question. You later see what people asked and did not get.
That third part gets overlooked and it is valuable. A list of unanswered questions is a list of what your website is missing.
Setting that rule is not a checkbox in an interface. It is a decision about how much the system may assume when a detail is incomplete. Companies that draw the line strictly get an assistant that hands off to a person more often. It almost never states something that is not true. Better to make that trade deliberately than to discover it after the first customer complaint.
Showing the source
A system that knows where its answer came from can show you. The answer arrives with a document name or a link to the page the detail came from.
In week one you read twenty conversations and see whether the system is reading what you think it is reading. Without sources, your only check is knowing every answer by heart.
ChatGPT versus an assistant tied to your content
This is the distinction most owners never see, because both look like a chat window. They behave nothing alike.
| Situation | Public tool (e.g. ChatGPT) | Assistant tied to your content |
|---|---|---|
| Source of answers | Anything the model has read | Only your approved content |
| When the detail is missing | Offers a plausible guess | Says it has no answer and hands off |
| Your prices and terms | Does not know them, infers from similar firms | Reads them from your price list |
| Verifiability | No source | Answer ties back to a document or page |
| When your terms change | Unaware of the change | Answers change with the content |
| Who carries the consequence | You, in front of the customer | Risk is bounded by what you approved |
Public tools are built for a different job. They are excellent for writing and thinking, and unsuitable for quoting your prices to your customers.
There is a middle option worth recognising. Some tools use your content but may top it up with general knowledge when a detail is missing. That sounds useful and is the most dangerous arrangement of the three, because the invented part looks identical to the verified part. Ask a vendor directly whether the system may say anything outside your documents.
Five tests to run before go-live
Every vendor will tell you their system does not make things up. These five questions check that in ten minutes, and you should ask them yourself, live.
- Ask about something that does not exist. Invent a service you do not offer and ask its price. The system must say it has no such service, not produce an estimate.
- Ask about a detail that appears twice on your site. If the two places disagree, a good system flags it or asks which you mean.
- Ask something outside your field. A clinic does not need a view on the weather forecast. The system must return to its subject.
- Ask for the source. Say "where did you get that." The answer should point to a specific page or document.
- Ask the same question three ways. The wording may differ, the substance may not. Different facts mean the system is guessing.
If it fails the first test, the rest do not matter. A confident answer to a question with no basis in your documents is the clearest possible sign that something is configured badly.
Testing does not end at the demo
Those five questions are a minimum, not the whole job. Before the system reaches customers, work through twenty questions you actually receive and compare the answers with what you would say yourself.
With our clients we run that round before go-live and again after the first week, once real customer questions come in. Repeat the same tests after any significant change to prices or terms, because a change in content changes what the system is entitled to claim.
What stays your responsibility
Restricting the system to your content solves guessing. It does not solve being wrong. If the source is wrong, the error gets relayed faithfully, in a confident voice.
- Contradictions in your content. The price on the page, in the PDF and in the brochure are often three different numbers. The system cannot know which is real.
- Outdated documents. Last year's price list still sitting on your site will be read as current.
- Rules written nowhere. If your cancellation terms live only in your team's heads, the system cannot honour them.
- Exceptions you make verbally. Discounts and side arrangements have to stay with people, because they exist in no document.
In our experience, pricing is the most common source of trouble. With several clients, the pricing on the website was both out of date and inconsistent between pages. The assistant gave wrong answers until we reconciled the sources. The system had not malfunctioned in those cases; it relayed exactly what the pages said.
Contradictions like that rarely surface on their own. A customer who was quoted the better terms has no reason to tell you something is off, so the error can sit there for months. The guide on preparing your content for an AI assistant walks through the sources step by step.
Restricting a system to verified sources and setting rules for unknown questions is work we do, along with testing answers before the system reaches customers. In more sensitive cases, such as an AI assistant in healthcare, we agree the line between answering and handing off with you before go-live.
Frequently asked questions
Short answers to the questions we get asked most.
Can you completely stop AI from making things up?
The risk drops to a very low level, but zero cannot be promised. A system restricted to your content, with a clear rule for unknown questions and monitored conversations, errs rarely and visibly.
How do I know my assistant is not guessing?
Ask it about a service you do not offer. If it answers with an estimate instead of saying it has no such information, the system is not properly restricted.
What happens when someone asks something our documents do not cover?
A properly configured system says it does not have that information and offers a contact route or a handoff. The question should also be logged, so you can see what is missing.
Is an AI agent riskier than a plain assistant?
Not on accuracy, but an agent acts, so a mistake carries further. That is why each action gets a defined limit on what may happen without human approval, covered in the guide on chatbot vs AI agent.
Who is liable if AI gives a customer wrong information?
You are, in front of the customer, exactly as with any employee. That is why what the system may assert, and what must go to a person, gets defined before go-live.
Accurate AI answers depend on configuration and content, not on luck or on picking a better model. A system restricted to your documents, and required to admit what it does not know, gets things wrong less often than an average new hire. The rest is on you, and it comes down to your content telling one story instead of three. That is worth doing whether or not you ever deploy AI.