Most people now use AI for decisions that matter, and most people still don't trust it.
That is not a contradiction. It is the actual, measured state of the market. Roughly 79% of companies now use AI in some part of their decision-making, yet only 46% of people globally say they trust AI systems. High adoption, low trust — running side by side, at scale, right now.
The instinct most people reach for to close that gap is "show me a citation." If an AI answer has a source attached, it feels safer. If it doesn't, it feels suspicious.
That instinct is understandable, and it is also the wrong test.
A citation tells you a source exists. It does not tell you whether the source actually supports the specific claim being made, whether it's the right kind of evidence for that claim, or whether "no citation" here actually means "no evidence" at all. Some of the most reliable claims an AI system can produce carry no external citation whatsoever — and some of the least reliable ones carry several.
The real question isn't "was this cited." It's "what is this actually standing on."
Claim
what is being asserted
Type
which of six evidence classes
Basis
is it actually grounded
Gap or artifact
tell the two apart
Verdict
trust it, or flag it
Why "no citation" stopped being a good enough test
The trust problem isn't abstract. It's measurable, and it's getting worse in specific, documented ways.
A 2025 study of AI-generated references found that 19.9% were completely fabricated, and a further 45.4% had serious bibliographic errors — meaning close to two-thirds of the citations an AI produced in that study were either invented or broken in some material way. Separately, tracking of fabricated citations in published research shows a roughly 12x increase between 2023 and 2025, with more than 4,000 biomedical papers found to cite research that doesn't exist.
The legal system is already absorbing the cost of this. In the year to February 2026, 66 US legal cases were sanctioned specifically for AI-generated errors, and by June 2026 that figure had grown to 1,598 court proceedings worldwide involving AI content with fabricated citations — with over $145,000 in AI-filing penalties in US courts in the first quarter of 2026 alone.
None of that is a reason to distrust AI outright. It's a reason to stop using "does it have a citation" as the whole test, because that test is too easy to pass without actually being true, and — just as importantly — too easy to fail while actually being correct.
There's a real complication on the other side of this too, worth stating honestly rather than skipping past: a Yale study found that people who used AI chatbots to help verify news saw a 15-point decline in their own unassisted ability to detect misinformation after just four weeks. Outsourcing verification isn't free. It can make you worse at doing it yourself. That's a genuine reason to want a system that shows its work rather than one that just hands you a clean-looking answer.
What a claim can actually be standing on
If a citation alone isn't the right test, the more useful question is: what type of evidence produced this claim? In practice, there are six real, distinct answers.
1. External research
A claim grounded in a live search or lookup against real, current sources — the kind of evidence you'd expect a citation to point to, and where a missing one is a genuine red flag.
2. User-supplied information
A claim built from something you told the system directly — your own brand facts, your own market knowledge, a correction you supplied. This doesn't need an external citation. The citation is you.
3. Evidence carried forward from an earlier step
In a multi-stage piece of research, a later stage can correctly rely on evidence a prior stage already established and verified, rather than re-researching it from scratch. The citation lives upstream, not on every downstream sentence that uses it.
4. Analytical inference, clearly labelled as inference
Sometimes the honest answer is "based on the evidence gathered, this is the most reasonable interpretation" — not a discovered fact, a judgement call. That's legitimate, provided it's labelled as inference and not dressed up as a citation.
5. Deterministic calculation
Some claims are just arithmetic. If the inputs are correct, a fixed, reproducible calculation doesn't need a citation any more than a spreadsheet formula does — the "evidence" is the fact that the same inputs always produce the same output.
6. Human correction, applied after the fact
This is the one most coverage of AI trust skips entirely: when a person with real, first-hand knowledge corrects an AI-generated claim, that correction is itself a form of evidence — arguably the strongest kind available, because it comes from someone who was actually there. We think of this as a real, sixth category in its own right, not just a quality-control step bolted onto the other five. That's a genuine synthesis on our part, not an established industry term — worth saying plainly rather than dressing it up as more than it is.
Six real answers. One test — "is there a citation" — was never going to cover all six.
The distinction almost nobody draws: a gap versus an artifact
Here's the angle that's genuinely underexplored, and it matters more than it sounds like it should.
When an AI system flags something as uncertain, that flag can mean one of two very different things:
A genuine evidentiary gap — the subject actually couldn't be verified. Nobody has published the number. The information isn't public. The uncertainty is real and about the world.
A language artifact — the underlying finding is fine, but it got flagged as negative or uncertain purely because of how it was phrased, not because of what it actually says. The uncertainty is about the sentence, not the subject.
These look identical on the page. Both show up as a yellow flag next to a claim. But they call for completely different responses. A real gap needs more research. A language artifact needs better phrasing — and treating it like a research gap wastes time chasing evidence for something that was never actually in doubt.
This distinction barely appears in existing coverage of AI trust and hallucination. Most of it stops at "AI can be wrong, verify important claims" — true, but not actionable. Knowing which kind of flag you're looking at is what actually tells you what to do next.
Why narrower questions beat broader ones
There's a second pattern worth naming, and it runs in the opposite direction from what most people assume "more thorough" AI research should look like.
The instinct is to ask one big, open question and expect a comprehensive answer back. In practice, the opposite tends to work better: breaking a broad question down into several narrow, specific, bounded ones — and asking them separately — produces evidence that's easier to check, easier to trust, and often more directly useful than one sprawling answer trying to cover everything at once.
Part of why: a narrow question gives a system less room to blur together things that should stay separate. It also gives you, the reader, a much easier claim to actually verify — a specific answer to a specific question is something you can check against a specific source. A single paragraph trying to answer five things at once usually can't be checked against anything in particular.
Less, aimed more precisely, tends to beat more.
What this actually looked like in real market research
We didn't take any of this on faith. We ran real research into how people currently describe their frustration with AI tools that make evidence and trust claims — the recurring language people actually reach for, not a hypothesis about it.
Three phrases came up often enough, independently, to be worth repeating here, without attaching them to any specific product or company, because the pattern is the point, not any one vendor:
- A felt "5% bullshit rate" — the sense that even a mostly-reliable AI tool has a small but real proportion of confidently wrong output baked in, and that this proportion rarely shrinks to zero however good the tool gets.
- The "illusion of transparency" — where a tool visibly shows its reasoning or sources, which reads as trustworthy, without that visible layer actually being checkable or correct.
- The "burden of validation" — the quiet, real cost that AI-assisted work doesn't remove the need to check things, it just moves the checking onto the human, later, often without warning that it's still required.
The same research also surfaced what people actually want instead, independently and repeatedly, without being prompted toward any particular answer: consistent, verifiable citation capability; a visible explanation of how the AI reached its conclusion rather than just the conclusion itself; and a synthesised, actionable verdict instead of an unfiltered dump of everything the system found.
We're not claiming credit for identifying that demand — the research found it on its own. What we'd point to is that it maps closely onto how we've built our own product: an honesty layer that shows assumptions, gaps and contradictions rather than hiding them; a way to correct and re-run an analysis with real context rather than starting over; and a single synthesised verdict up front, with the full evidence trail available underneath it for anyone who wants to check our working.
A checklist for the next AI claim you're handed
Next time an AI system — ours or anyone else's — hands you a claim you're about to act on, five questions get you further than "is there a citation":
- What type is this? External research, your own input, carried-forward evidence, labelled inference, a calculation, or a human correction — they call for different levels of trust.
- If it's a citation, does the source actually say this? A citation existing and a citation supporting the specific claim are not the same check.
- If it's flagged uncertain, is that a gap or an artifact? Ask what specifically is unverified — the subject, or just the sentence.
- Would a narrower version of this question get you a more checkable answer? If the claim is broad, consider whether splitting it would help.
- Has anyone with first-hand knowledge actually looked at this? A human correction, where one is available, is often the strongest evidence type on the list — use it.
None of that requires distrusting AI outright. It requires treating "cited" and "true" as two different questions, because they are.
The trust gap in AI isn't going to close by adding more citations to more sentences. It closes by being honest about what kind of evidence a claim actually rests on — and by making the difference between "we don't know yet" and "this just reads badly" visible, instead of collapsing both into the same yellow flag.
Frequently Asked Questions
Not necessarily. A claim can be legitimately uncited if it comes from information you supplied yourself, a deterministic calculation, a clearly labelled inference, or evidence already verified earlier in the same piece of research. The test is what type of evidence it rests on, not whether a citation is attached.
A genuine gap means the underlying subject actually couldn’t be verified — the information isn’t available. A language artifact means the finding itself is fine, but it was phrased in a way that reads as negative or uncertain. Both can show up as the same warning flag, but a gap needs more research while an artifact needs better wording.
A narrow, specific question gives a system less room to blur separate ideas together, and gives you a specific, checkable claim to verify against a specific source. A single broad answer trying to cover several things at once is much harder to check against anything in particular.