Why ChatGPT makes things up, explained in plain words

A chip motif standing for how a chatbot produces its answers

In June 2025 the High Court in London dealt with two cases where lawyers had put made-up law in front of a judge. In one of them, a list of 45 citations turned out to include 18 cases that did not exist at all. Many of the real ones did not say what they were quoted as saying.

The judges spelled out why this happens with freely available AI tools such as ChatGPT: “The responses may make confident assertions that are simply untrue. They may cite sources that do not exist.”

That is the odd part, and it is worth sitting with. A machine that can write a tidy legal argument will also invent a court case to put in it, in the same calm tone, without a flicker. Why doesn’t it just say it doesn’t know?

It learned what answers look like

A chatbot starts life by reading an enormous amount of text and learning its patterns: which words go with which, what a sentence about a court case looks like, how a date is usually written. Where there is a strong pattern, it gets very good. OpenAI’s own researchers point out, in a 2025 paper on exactly this question, that these models rarely make spelling mistakes, because spelling follows consistent patterns.

Some facts have no pattern, though. Nothing about a person’s name tells you their birthday. The paper’s authors tested this on one of themselves: asked for his birthday and told to answer only if it knew, a leading open-source model gave three different wrong dates on three attempts.

Here is the mechanism in one line. A fact the model saw many times, it tends to get right (the paper uses Einstein’s birthday as the example). A fact it saw once or never, it can only fill in with something that looks right. The authors put a rough number on it: if 20% of birthday facts appear exactly once in the training data, expect the model to get at least 20% of birthdays wrong.

So the model is not lying. It has no separate list of true things to check against. It has a very good sense of what an answer should look like, and a date-shaped answer looks just like a date.

[[photo:1]]

The exam where a blank scores zero

That explains where the wrong answers come from. It does not explain why the model states them so confidently instead of hedging. For that, the OpenAI paper offers an analogy that holds up well.

Think of a multiple choice test that gives one mark for a right answer and nothing for a blank. If you do not know, you should always guess. A guess sometimes scores; a blank never does. Over a whole paper, the student who bluffs beats the student who leaves honest gaps.

The researchers argue that AI models are graded the same way. Many of the tests used to compare them score right or wrong, with no credit for “I don’t know”. A model tuned to do well on those tests learns to bluff. And, as the paper says of students, bluffs tend to be “overconfident and specific, such as ‘September 30’ rather than ‘Sometime in autumn’”. A specific, confident wrong answer is exactly what the lawyers got.

Their proposed fix is the one some exams already use: penalise a wrong answer more than a blank, and say so in the instructions. The paper names India’s JEE and NEET entrance exams among the tests that have done this. Whether the big AI leaderboards follow is still an open question.

Europe went and counted

This is not only a problem for obscure birthdays. In October 2025 the European Broadcasting Union and the BBC published a study in which journalists from 22 public broadcasters in 18 countries, working in 14 languages, checked more than 3,000 answers from ChatGPT, Copilot, Gemini and Perplexity to questions about the news.

Almost half the answers, 45%, had at least one significant issue. A fifth contained major accuracy problems, “such as hallucinated and/or outdated information”. The biggest single problem was sourcing: 31% of answers had serious sourcing issues, including claims that the linked source did not actually support.

One example makes the point better than the percentages. Asked “Who is the Pope?”, ChatGPT, Copilot and Gemini each named Pope Francis in some answers, though he had died in April 2025 and Leo XIV had succeeded him. One Copilot answer named Francis as Pope while also stating the date he died.

There is some good news in the same report. Comparing the BBC’s own questions across two rounds, the share of answers with significant issues fell from 51% to 37%. The tools are getting better. They are not yet at the point of being taken on trust.

[[photo:2]]

What this means when you use it

Once you see the model as a very well read guesser, you can predict where it will slip. It is at its best on things that have been written about thousands of times: explaining photosynthesis, tidying a paragraph, summarising a document you pasted in yourself. It is at its weakest on the rare and the specific: a name, a date, a figure, a quote, a page number, a citation, and anything that changed recently.

That last one matters for a reader in Singapore as much as one in London. A question about a school policy, a fee or a rule that changed this year is exactly the outdated-information problem the EBU study kept finding, and those answers arrive in the same confident voice as everything else.

So the habit is small. When an answer hands you a specific detail you plan to rely on, go to the original: the official page, the actual document, the real report. The High Court told lawyers much the same thing, pointing them to official databases of legislation and judgments. If the chatbot gives a link, open it and check that it says what the chatbot claims, since that is where the European study found the most trouble.

Children can learn this habit too, and we wrote about teaching kids to fact-check an AI separately. The idea underneath is the same for everyone. A fluent answer tells you the model knows what an answer looks like. Only the source tells you it is true.

Was this useful?

Sources

  1. Ayinde v London Borough of Haringey and Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin) · Courts and Tribunals Judiciary (England and Wales)
  2. News Integrity in AI Assistants · European Broadcasting Union and BBC
  3. News Integrity in AI Assistants (report page) · European Broadcasting Union
  4. Why Language Models Hallucinate · OpenAI
  5. Why Language Models Hallucinate (arXiv:2509.04664) · arXiv