Blog

New working paper – Accuracy Paradox: Addressing Epistemic, Manipulative, and Societal Risks of Hallucination in AI Governance

Posted on    by
BlogWorking papers

New working paper – Accuracy Paradox: Addressing Epistemic, Manipulative, and Societal Risks of Hallucination in AI Governance

By 8 May 2026No Comments

CREATe is happy to present a new entry in our working paper series: Accuracy Paradox: Addressing Epistemic, Manipulative, and Societal Risks of Hallucination in AI Governance. This paper is co-authored by Zihao Li (CREATe, University of Glasgow), Weiwei Yi (CREATe, University of Glasgow), and Jiahong Chen (University of Sheffield). It builds on the authors’ earlier work, Zihao Li “Why the European AI Act transparency obligation is insufficient” (2023) in Nature Machine Intelligence, and the CREATe response to the UK Information Commissioner’s Office (ICO) consultation on the accuracy of training data and model outputs in generative AI.

This is the pre-print version of an article forthcoming in Computer Law and Security Review.

The rapid deployment of generative AI, and in particular large language models (LLMs), has transformed how information is produced, circulated and consumed in domains ranging from healthcare and education to law. Alongside this transformation, an urgent concern has emerged: AI hallucination. Several examples already illustrate its practical harms. A Norwegian user was told by ChatGPT that he had murdered his children, an American law professor was falsely named in a fabricated list of scholars accused of sexual harassment, complete with a citation to a non-existent Washington Post article, and Air Canada was ordered to honour a refund policy invented by its own chatbot. Such incidents have fuelled a regulatory consensus that the answer lies in making AI systems more accurate.

Across the UK ICO, the European Data Protection Board, the European Data Protection Supervisor, national DPAs in Belgium and Sweden, and leading industry players such as OpenAI and Google, accuracy has become the dominant metric for measuring and mitigating hallucination. Tech firms routinely use accuracy benchmarks as proxies for hallucination reduction, and a recent evidence review of legal LLM studies found that over 80% rely on quantitative accuracy-style measures. Accuracy, in short, has become the unchallenged north star of trustworthy AI.

This working paper argues that this accuracy-centric approach is conceptually narrow and normatively insufficient. The paper conceptualises the accuracy paradox: the very efforts aimed at reducing hallucination through hyper-optimisation of accuracy can paradoxically exacerbate existing harms or give rise to new ones. Overreliance on accuracy risks creating a false sense of epistemic certainty, amplifying users’ blind trust, and weakening the critical checks and balances on which responsible AI use depends.

The paper first develops a taxonomy of hallucination situated within a sociotechnical context, showing that hallucination is far more varied than factual error alone. Beyond factuality hallucinations, the taxonomy captures consistency hallucinations, reference fabrication, sycophancy, consensus illusion, oversimplification, and prompt-sensitivity hallucinations such as sandbagging and emotionally induced drift. This heterogeneity already suggests that a narrowly defined concept of accuracy cannot plausibly guide regulation on its own.

Click on image to download

The paper first develops a taxonomy of hallucination situated within a sociotechnical context, showing that hallucination is far more varied than factual error alone. Beyond factuality hallucinations, the taxonomy captures consistency hallucinations, reference fabrication, sycophancy, consensus illusion, oversimplification, and prompt-sensitivity hallucinations such as sandbagging and emotionally induced drift. This heterogeneity already suggests that a narrowly defined concept of accuracy cannot plausibly guide regulation on its own.

Building on this taxonomy, the paper then conceptualises the accuracy paradox along three dimensions:

  • Accuracy vs. Trustworthiness – Statistical accuracy is not the same as truth and trustworthiness. A model can become better at predicting what sounds correct without any grounding in whether the content is actually justified or verifiable. Overreliance on accuracy can create over-trust, erode meaningful transparency and interpretability, and mask epistemic fragility behind the appearance of reliability.
  • Accuracy vs. Autonomy – Optimising for surface-level fluency and rhetorical confidence can facilitate manipulation through outputs that are “not inaccurate” yet subtly persuasive, from LLM-powered advertising embedded in ordinary conversation to personalised nudging, sycophantic alignment, and even emergent forms of deceptive alignment that elude benchmark-driven evaluation.
  • Accuracy vs. Social Progression – At the collective level, accuracy-centric optimisation risks entrenching equity harms (social sorting, re-identification), plurality harms (epistemic convergence, compression of minority and non-Western perspectives), and criticality harms (cognitive deskilling, loss of moral attention, homogenisation of creative output).

The paper then turns to EU law and asks how well existing frameworks are equipped to confront the accuracy paradox. Through a close examination of the EU AI Act, the GDPR, and the DSA, it finds that accuracy is either treated as a high-risk-only obligation (AI Act Art. 15), operationalised through deterministic, record-based notions of data correctness (GDPR Art. 5(1)(d) and Art. 16), or framed as a procedural, diagnostic metric for platform audits (DSA Art. 15(e) and Recital 96). None of these instruments adequately addresses harms that arise precisely from outputs being fluent, aligned, and technically not inaccurate. Manipulation prohibitions in AI Act Art. 5(1)(a) require a high bar of intent and significant harm that emergent LLM behaviours rarely meet; GDPR Art. 22 remains tied to individual decisional outcomes rather than diffuse epistemic influence; and the DSA leaves “accuracy-as-disguise” largely outside its regulatory scope.

Finally, the paper sketches future directions that move beyond accuracy. These include reorienting LLM development towards epistemic trustworthiness (calibrated confidence, uncertainty signalling, justification, abstention), operationalising this through layered epistemic impact assessments integrated with Fundamental Rights Impact Assessments and systemic risk mitigation; embracing pluralism by design so that models surface diverse sources and contested perspectives; and reassessing hallucination itself as a context-sensitive phenomenon that can, in exploratory settings, support rather than undermine human thinking.

Accuracy Paradox: Addressing Epistemic, Manipulative, and Societal Risks of Hallucination in AI Governance

Zihao Li, Weiwei Yi, Jiahong Chen

CREATe Working Paper 2026/02

Abstract

The rise of generative AI has intensified concerns around AI hallucination, which involves outputs that are fabricated, misleading, oversimplified or untrustworthy. While many technical and policy responses treat hallucination as a failure of factual accuracy, this paper argues that such a narrow lens underestimates the complexity of the problem. AI hallucination is not merely a matter of truth or falsehood, but a multifaceted phenomenon with cognitive, communicative, and societal implications. Overreliance on accuracy has counterproductive effect: the accuracy paradox. We propose a taxonomy and theoretical framework for understanding hallucination risks across three dimensions: epistemic reliability, Human-AI interactive influence, and social impact. Through regulatory analysis, we show that accuracy-driven approaches often overlook harms such as illusion of consensus, subtly persuasive misinformation, and diminished social progression. Current legal regulation, including the EU AI Act, GDPR, and DSA, struggle to address these subtler forms of distortion. We call for regulatory strategies that go beyond static verification, embracing pluralistic, context-aware, and manipulation-resilient approaches to AI trustworthy governance.

Full paper can be downloaded here.