AI Hallucinations in UK Courts: The Cases, the Consequences, and How to Prevent Them
There is now a public database that does nothing but catalogue the moments generative AI went wrong in court. Maintained by the legal researcher Damien Charlotin, the AI Hallucinations Cases database tracks every documented instance where a court has found, or clearly implied, that a party relied on fabricated, AI-generated legal material. At the time of writing it records more than 1,700 cases worldwide, and the UK section alone runs to over fifty judgments spanning thirty different courts and tribunals, from the First-tier Tribunal all the way to the Court of Appeal.
This is no longer a novelty. It is a recognised, sanctionable failure mode of AI-assisted legal work, and the pattern behind it is consistent enough that it can be caught before a document is ever filed. This article walks through what UK courts have actually seen, what it has cost the people responsible, and the specific verification steps that prevent it.
Three ways an AI invents the law
Read enough of these judgments and the hallucinations sort themselves into three recurring types. The database uses roughly the same taxonomy, and each one fails a different verification check.
1. The case that never existed
The most obvious failure is a citation that resolves to nothing at all. The model produces a plausible case name, a real-looking court code and a sequential number, and none of it corresponds to a real judgment.
The landmark UK authority here is Ayinde v Haringey and Al-Haroun v Qatar National Bank, decided together by the Divisional Court on 6 June 2025. Between the two matters, a judicial assistant identified eighteen non-existent authorities across the witness statements and submissions, including a fabricated authority attributed to the very judge hearing the case. The court dealt with it under its Hamid jurisdiction, the mechanism for enforcing the duties lawyers owe to the court. It declined to make a finding of contempt on the facts but referred the lawyers to their professional regulators, and used the occasion to issue general guidance warning that filing fabricated material can amount to contempt of court or even perverting the course of justice.
It keeps happening. In one 2026 First-tier Tribunal appeal, a litigant relied on three cases named "Harrison", "Baxter" and "Hicks". The Tribunal and HMRC checked, found none of them existed, and the litigant admitted to using AI. This is the type a simple existence check catches every time: does a real judgment sit at that citation, or not?
2. The real citation bolted to the wrong case
This is the dangerous one, because a careless check passes it. The neutral citation is genuine (you can look it up and find a real judgment) but it has been attached to the wrong case name, or cited for a proposition it does not support.
In Ndaryiyumvire v Birmingham City University (County Court, October 2025), a legal representative cited "Z Ltd v A Ltd [2011] EWCA Civ 110". The citation is real; it just belongs to an entirely different, unrelated judgment. The court made a wasted costs order and referred the matter to the Bar Standards Board. In Mrs Sabrena Rodney v Gee'z Micro Bar (County Court at Dudley, April 2026), Grace v Black Horse Limited was cited as "[2014] EWCA Civ 1091", but that neutral citation actually relates to R (Grace) v Secretary of State for the Home Department, a completely different case. That one drew an SRA referral.
Verifying that "the citation exists" is not enough here. The only thing that catches this class of error is checking that the resolved judgment's real name matches the name you were given.
3. The quote that was never written, and the statute that doesn't exist
The third type keeps the case but invents its contents. In PSAHSC v Tchampet (High Court, January 2026), written submissions produced via Microsoft Copilot included a purported quotation from paragraph 21 of Gupta v General Medical Council [2001] UKPC 61, a real case, but the quoted passage simply was not there. A closely related variant fabricates the statute itself. Across several tribunal cases, litigants have cited non-existent statutory provisions and rules: a fictitious section of the Housing Act, an "Interest on Debts (Scotland) Act" extract the court found did not exist, an invented insolvency rule "expressly providing" for an order it does not mention.
Catching this means going past the citation to the judgment's actual text: is the quoted paragraph really there, and does the case really say what it is cited for?
The consequences are not theoretical
None of this ends with a polite word from the bench. The UK judgments in the database carry real costs, financial, professional and reputational.
| Case | Court | What the AI produced | Consequence |
|---|---|---|---|
| Ayinde v Haringey; Al-Haroun v QNB | Divisional Court | 18 non-existent authorities | Referral to SRA / Bar Standards Board; contempt guidance |
| Ndaryiyumvire v Birmingham City Univ. | County Court | Real citation, wrong case | Wasted costs order; Bar referral |
| Mrs Sabrena Rodney v Gee'z Micro Bar | County Court, Dudley | Misattributed neutral citation | SRA referral |
| The Father v The Mother | High Court | Fabricated authorities (pro se) | £5,900 costs awarded against him |
| Mr J Harrison v Leeds Gymnastics | Employment Tribunal | Fabricated case law | £2,178 preparation-time order |
| Holloway v Beckles | First-tier Tribunal | Fabricated case law | £750 costs order |
Beyond costs orders and regulatory referrals, courts have refused strike-out applications for "unreasonable conduct", demanded signed statements of truth about AI use, and, in the Ayinde guidance, put the profession on notice that the next lawyer may not escape a contempt finding. Notably, the problem is not confined to litigants in person. Of the UK cases, a substantial share involve regulated lawyers, and the tools named include not just ChatGPT but Microsoft Copilot, Google's AI Overview, and AI features built into legal-research software.
Why AI fabricates cases in the first place
It helps to understand why this happens, because the reason explains why it will keep happening until something checks the output. A large language model is not a database and does not "look up" the law. It is a next-token predictor: given the text so far, it generates the most statistically plausible continuation, one word-piece at a time. It is optimised to produce text that looks right, and it has no internal concept of whether a given statement is true.
Legal citations are the perfect trap for this kind of system, for three compounding reasons.
The format is rigidly patterned. A citation like [2021] EWHC 2556 (Ch) follows a strict, learnable template: bracketed year, court code, running number, division. The model has seen millions of real ones, so it can generate a new string in exactly the right shape effortlessly. The problem is that a well-formed citation and a real citation are indistinguishable on the surface. The model produces fluent structure with no guarantee that a judgment actually sits at that coordinate.
Plausible-but-wrong is the default, not the exception. Because the model is completing a pattern rather than retrieving a fact, "a case that sounds like it should exist" is exactly what it produces when the real one is missing from its training data or its memory of it is fuzzy. The most-studied example, Mata v Avianca in the United States, is instructive: the tool produced six citations with fake names, fake docket numbers and fake reasoning, and when the lawyer asked it to confirm the cases were real, it insisted they could be found on Westlaw and LexisNexis. The model cannot tell the difference between recalling a case and inventing one, so it cannot reliably flag its own mistakes.
Confident guessing is rewarded. The way these models are trained and evaluated tends to punish "I don't know" and reward a confident, well-formed answer, so under uncertainty the model guesses rather than abstains. That is why the output is always delivered in the same authoritative tone whether it is right or wrong. The style carries no signal about reliability.
The scale of the effect is well documented. A 2024 Stanford study (Dahl et al.) found general-purpose models hallucinated on the majority of specific legal queries, and follow-up work found that even purpose-built legal research tools produced incorrect or fabricated information in a meaningful share of answers. In other words, this is not a bug in one chatbot. It is a structural property of how the technology works, which is precisely why every AI-generated citation needs an external check before anyone relies on it.
Why a manual check isn't enough
Knowing AI hallucinates, you might think careful reading catches it. In practice, manual verification fails for two reasons. The first is volume: in one 2025 Employment Tribunal case, a single set of submissions contained forty-six citations, nine wholly fictitious and thirty-seven misrepresenting what the real case said. Checking each one by hand, against the right source, is exactly the work busy practitioners skip under deadline. The second is the wrong-name trap from category two: the lazy check ("does this citation exist?") returns yes, and the reader moves on, never noticing the name attached to it is wrong.
How CaseNode stops it before filing
CaseNode is built around the three checks that map exactly onto the three failure modes above. Paste a paragraph, a skeleton argument, or a full document into the Validate tool and every citation is checked against the real UK corpus at once.

Existence check, catches the case that never existed. Each citation is resolved against the corpus. Anything that does not correspond to a real judgment comes back flagged, like the "Not found" result above. If a citation cannot be resolved, it is fabricated, and you stop before it reaches a judge.
Name match, catches the real citation on the wrong case. When a citation does resolve, CaseNode surfaces the true case name sitting at that coordinate, so a mismatch with the name in your document is obvious. This is the check that would have caught the Ndaryiyumvire and Rodney misattributions, the ones a plain existence check waves through.
Read the source, catches the invented quote and the misapplied authority. Because every judgment in the corpus is fully searchable down to the paragraph, you can jump straight to the paragraph a quote is attributed to and confirm the words are actually there.
CaseNode also shows you, in one place, every recent case that has cited the authority and how it treated it, whether it was followed, distinguished, doubted or overruled. That is how you check the AI applied the case correctly, not just that it exists. If a submission cites a judgment for a proposition, you can see at a glance whether later courts actually read it that way, and whether it is still good law. This is the check that would have caught the many entries in the database where a real case was cited for something it never decided.

For teams building AI into their own products, the same checks are available as a single API call, so a citation can be verified the moment it is generated rather than after it is filed:
curl https://api.casenode.ai/validate \
-H "X-API-Key: $CASENODE_API_KEY" \
-G --data-urlencode "text=See Z Ltd v A Ltd [2011] EWCA Civ 110."
{
"citations": [
{
"cite": "[2011] EWCA Civ 110",
"status": "name_mismatch",
"resolved_case": "TS (Burma) v Secretary of State for the Home Department",
"given_name": "Z Ltd v A Ltd"
}
]
}
A not_found status is your fabricated-citation signal; a name_mismatch is the far more dangerous misattribution. Either way, the check happens before anyone relies on the citation.
Better still: stop the hallucination at the source
Everything above catches a hallucination after the AI has produced it. The stronger fix is to stop it happening in the first place, and that comes straight from the mechanism explained earlier. A model hallucinates because it is generating citations from statistical memory rather than reading real cases. Give it the real cases to work from, and the reason to invent them disappears.
That is the other way to use CaseNode. Through its API, its MCP server and its research skills, an AI assistant can search the real UK corpus and pull actual judgments into its context before it writes a single word. Instead of asking a model "what cases support this argument?" and hoping its memory is accurate, the assistant retrieves genuine, verified cases from CaseNode, reads the relevant paragraphs, and drafts from that grounded material. The citations it produces are real because they were retrieved, not recalled.
This is the difference between a model working from memory and a model working from the source. When the AI is grounded in real cases, with the correct context in front of it, it is reasoning over fact rather than filling a plausible-looking gap. The verification checks above remain a sensible final safeguard, but grounding the research in CaseNode from the outset removes the condition that produces hallucinations to begin with.
The takeaway
Fifty-plus UK judgments, thirty courts, real costs orders and regulatory referrals: AI hallucinations in legal work are now a documented, recurring problem with a well-understood shape. Every case in the database failed one of three checks. The citation did not exist, the citation belonged to a different case, or the case did not say what it was cited for.
Those three checks are mechanical, and machines are better at them than tired humans reading forty-six citations at midnight. Used properly, generative AI is a genuine accelerant for legal research. The difference between AI-assisted work you can file and AI-assisted work that gets you referred to the SRA is a verification step, and CaseNode makes it a single query.