
“Make Me a Bibliography”: What AI Gets Wrong and How to Get a Real One
Ask a chatbot for a bibliography and you get perfect formatting around entries that may not exist — here is what the measured error rates are and how to get a list that survives a check.
Type "make me a bibliography on X" into ChatGPT or Claude and you will get a list back in seconds, formatted correctly, in the style you asked for, with volume numbers and page ranges that look exactly like the real thing. The formatting is usually right. The existence of the entries is the part nobody guarantees. A model that has never queried a bibliographic database is assembling plausible citation-shaped text from patterns in its training data, which means some entries will be real, some will be real papers with the wrong year or the wrong first author, and some will be nothing at all.
That distinction matters because a reference list is not prose. If a paragraph is a bit vague, a reader shrugs. If a reference does not exist, a supervisor or reviewer can establish that in thirty seconds with a search box, and checking has become routine.
What the model is actually doing with that prompt
Generative models do not look references up unless a tool is wired in to do it. Jin and colleagues describe the mechanism plainly in their 2026 review of reference management software: AI-generated citations reproduce author names, article titles and publication details that mimic the structure of real citations, rather than retrieving verifiable sources (p. 2). The output is shaped like a bibliography because bibliographies are highly patterned, and patterns are what these systems are good at.
Resnik and Hosseini, writing in 2026 on whether hallucinated citations amount to research misconduct, make the same point about why this does not simply get fixed: the problem is "inextricably linked to how LLMs operate" (p. 2). Validation layers have lowered the rate for some tools, but the underlying generation step is still generation, not retrieval.
How often the entries are wrong
The measurements are not reassuring, and they vary enormously by model and field.
Chelli and colleagues (2024) ran eleven systematic reviews as prompts through three models and analysed 471 returned references. They counted a paper as hallucinated if any two of title, first author or publication year were wrong. Hallucination rates came out at 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Bard. Precision, meaning the share of returned papers that were actually relevant to the review, was 13.4% at best (p. 1). Their conclusion was that these tools should not be the sole or primary means of assembling a literature base, and that anything they return needs validating by the author (p. 9).
More recent numbers are lower but still material. Misra and Udandarao (2026) issued literature-style prompts to three frontier models and verified the extracted citations against Crossref, OpenAlex and Semantic Scholar. They report roughly 12% hallucinated for GPT-4o, 16% for Claude 3.5 Sonnet and 21% for LLaMA-4 Maverick. They also note published rates as high as 88% in legal contexts, and that on the CiteME benchmark models scored 4.2–18.5% accuracy against 69.7% for human annotators.
Resnik and Hosseini collect the consequences. A paper in Academic Ethics had 19 of 29 citations hallucinated; a study retracted by PLOS ONE had 18 of 76; a piece in Digestive Diseases and Sciences had 12 of 14. A study of GPT-4o citations in mental health found 56% contained errors, with one in five fabricated outright (p. 3). Their argument goes further than embarrassment: where citations function as data supporting a finding, they hold that fabricated ones can constitute provable research misconduct, and that using a tool without checking its output counts as recklessness rather than bad luck (p. 7).
Three different problems, one prompt
"Fake" is too blunt a word for what comes back. Misra and Udandarao label their benchmark entries in three categories, and the distinction is the practical one for anyone cleaning up a generated list:
- Valid — all major fields correct. The paper exists and you can open it.
- Partially valid — some overlap, one or more errors. A real paper with a mangled year, a co-author promoted to first author, a title that drifted. These are the dangerous ones, because a quick glance passes them and a reviewer's search does not.
- Hallucinated — no credible match found. The entry is an artefact.
There is a fourth case their labels do not cover: the reference is entirely real and entirely correct, and says nothing that supports the sentence you attached it to. Formatting checks never catch that one.
Before and after
Before. Prompt: "Make me a bibliography of 10 sources on citation hallucination in APA 7." Output: ten APA-shaped entries, DOIs included. You paste them in. Two do not resolve. One is a real article with the wrong journal. One is a conference paper from a year the conference did not run. You find this out when your supervisor pastes a title into Google Scholar during the meeting.
After. Same topic, different order of operations. You find the sources first, in something that queries an index, then you format. Every entry that lands in the list has a DOI or a catalogue record behind it, you have opened at least the abstract, and the page number you cite came from the document rather than from a model's guess. The list will usually be shorter than the generated one, and every row on it holds up when someone searches for it.
The order matters more than the tool. Retrieve, verify, then format; never format first and hope.
Getting a bibliography that survives a check
A reference manager alone will not save you here, and it is worth being clear why. Jin and colleagues point out that market-leading systems were built to organise, store and format references, not to verify that they exist. EndNote's "Find Full Text" retrieves the full text of a reference you already have, but it does not confirm that the citation metadata correspond to a real publication (p. 5). Subramanian and colleagues (2025), surveying citation practice, describe reference managers in exactly those terms: collecting, storing, organising, and generating a bibliography in the preferred style (p. 97). Style compliance and existence are separate properties, and the older tools were designed for the first one.
What actually closes the gap is a lookup against an authoritative index. Misra and Udandarao's pipeline is instructive because it shows how much a plain lookup gets you: exact matching of title plus first author against bibliographic databases caught around 65% of hallucinations on its own, fuzzy retrieval pushed recall to about 75%, and adding a verification step reached roughly 80% precision. You can do the cheap version by hand for ten references. Paste the title into Crossref or the DOI into doi.org, and check that the first author and year match what you were given.
If you would rather not do that by hand, that is what cytado's bibliography generator is for: entries come out of a corpus built from Crossref, OpenAlex, DOAJ and library indexes, so the thing you are formatting has already been confirmed to exist. If you already have a list from a chat window and just want to know which rows are real, run it through the source checker instead and keep the survivors. Every academic source cited in this article was retrieved that way, which is also why there are five of them rather than fifteen.
Two related reads if you are working this way: does AI invent sources, and how to verify a citation, and how to check if a source really exists.
Frequently asked
- Can ChatGPT make a bibliography that is actually correct?
- It can format one correctly. Whether the entries exist is a separate question, and it depends on whether the model is retrieving from an index or generating from patterns. Measured hallucination rates in published studies range from roughly 12% for recent models on general prompts to over 90% in one 2024 comparison, and up to 88% in legal contexts. Treat any unverified generated list as a set of candidates, not a bibliography.
- How do I check a bibliography the AI gave me?
- Take each entry and match title plus first author against a bibliographic index such as Crossref or OpenAlex, and resolve any DOI through doi.org. In one published detection pipeline, exact matching alone caught around 65% of hallucinated citations. Pay particular attention to entries that nearly match a real paper, since a wrong year or a promoted co-author passes a casual glance.
- Will Zotero or EndNote catch a fake reference?
- No. Reference managers were built to organise, store and format references, not to verify that they exist. EndNote's full-text lookup retrieves a document for a reference you already have; it does not confirm that the metadata correspond to a real publication. Verification needs a query against an authoritative index.
- Is citing a hallucinated source actually a serious problem?
- It can be. Beyond retractions — documented cases include papers with 19 of 29 and 12 of 14 references fabricated — one 2026 analysis argues that where citations function as data supporting a finding, fabricated ones can amount to provable research misconduct, and that skipping verification counts as recklessness rather than an honest mistake.
Sources
Read next
- Annotated bibliography checker: verify every source before you annotate itCheck that every entry on your annotated bibliography resolves to a real publication before you spend 150 words annotating it, and learn why reference managers and self-checking AI will not catch fabricated references.
- How to Use AI for Finding Sources Without Getting BurnedAI is safe for mapping a field and unsafe for writing your reference list; here is the line between the two and the checks that turn an AI-assisted search into a bibliography you can defend.
- Will Your Supervisor Spot AI-Generated Sources?AI writing detectors are unreliable, but a fabricated reference either resolves in a database or it does not, which is exactly why fake sources are the part of AI-assisted work that gets caught.