
How to Use AI for Finding Sources Without Getting Burned
AI is safe for mapping a field and unsafe for writing your reference list; here is the line between the two and the checks that turn an AI-assisted search into a bibliography you can defend.
You can use AI to find sources for an essay or a thesis, just not in the way most students do it. Ask a chatbot for ten references on your topic and a sizeable share of what comes back will not exist, while a good part of the rest will carry the wrong year, the wrong volume or a page range that leads nowhere. The safe version of this workflow keeps the model in the job it is genuinely good at, mapping a field and giving you the vocabulary the literature actually uses, and hands the reference list over to a database that can confirm each item is real. Below is where that line runs, what goes wrong when you cross it, and the sequence of checks that turns an AI-assisted search into a bibliography you can defend.
Three ways an AI-sourced citation fails
The paper does not exist. In a study of 636 references generated for 42 multidisciplinary topics, 55% of the GPT-3.5 citations and 18% of the GPT-4 citations were fabricated, meaning they did not correspond to any real scholarly work (Walters and Wilder, p. 1). Across six studies pooled by the same authors, 51% of 732 citations were fabricated (p. 2). A comparison focused on systematic reviews found hallucination rates of 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Bard, where a reference counted as hallucinated if any two of title, first author or year were wrong (Chelli et al., p. 1).
The paper exists, but the details are wrong. Among the GPT-3.5 references that pointed at real works, 43% still contained substantive errors, against 24% for GPT-4 (Walters and Wilder, p. 1). More than a third had an incorrect volume, issue or page number and 22% gave the wrong year, often because an online posting date replaced the original date of publication (p. 4). Those are precisely the fields a supervisor spot-checks, and they are the fields a model gets wrong most often.
The paper exists, the details are right, and it does not support your sentence. In the systematic-review comparison, precision was 9.4% for GPT-3.5, 13.4% for GPT-4 and 0% for Bard, so even the references that turned out to be real were mostly not relevant to the question being reviewed (Chelli et al., p. 1). A real reference hung on a claim it does not make will still fail the check, because the page your supervisor opens will not contain the sentence you attributed to it.
Here is the difference in practice.
Before, pasted straight out of a chat window:
Smith, J. A. (2019). Motivational factors in second language acquisition. Journal of Applied Linguistics, 40(3), 215–233.
Everything about it looks right. The journal sounds plausible, the volume and page range are in a normal shape, the year is recent enough. None of that is evidence of anything, because a fabricated reference is built from exactly those patterns.
After, once it has been through a database:
The DOI resolves to a landing page with that title. The author's surname and initials match the publisher record. The volume, issue and page range match too. You have opened the PDF and found, on p. 219, the sentence that actually carries the claim you attached to it. Now it is a citation.
Why a better prompt does not solve this
A language model produces text by predicting which token is most likely to follow the tokens before it (Banh and Strobel, p. 4). A bibliographic reference is one of the most predictable strings in academic writing: surname, initials, year in brackets, title, journal, volume, pages. A model can assemble one that is perfectly shaped without any record standing behind it, which is why fabricated references look legitimate at first glance (Walters and Wilder, p. 2).
This is not a bug waiting for the next release either. Farquhar and colleagues describe a subset of hallucinations as confabulations, arbitrary and incorrect generations, and their detection method, which measures uncertainty at the level of meaning rather than wording, explicitly does not cover cases where a model is confidently wrong for systematic reasons (p. 629). Detection is an active research problem, not a solved one. A 2026 review of the field puts reported citation hallucination rates between 18% for GPT-4 and over 70% for other frontier models, rising to 88% in legal contexts, and notes that on the CiteME benchmark models score 4.2–18.5% where human annotators reach 69.7% (Misra and Udandarao). Rates differ by model and by field, and no version of "use the newer model" gets you to zero.
Does asking the chatbot to check its own references help?
Partly, and not enough to stop there. In one experiment, ChatGPT generated 40 references with DOIs for musculoskeletal rehabilitation topics; only 7.5% were fully accurate on first pass and 42.5% were completely fabricated. After the model was asked to verify and revise its own references, the share of fully accurate ones rose to 77.5% (Safran and Çalı, p. 695). That is a genuine improvement, and it still leaves roughly one reference in five wrong, with at least one item in the first prompt set still fabricated after the check (p. 697). Walters and Wilder add the awkward part: chatbots often give incorrect answers when asked directly whether a citation is correct or whether they fabricate citations (p. 2).
Treat a self-check as a way of thinning the list, never as verification. The thing you need is an external record.
A workflow that survives your supervisor
- Use the model for orientation. Ask it to explain how a debate is structured, which schools of thought argue what, and which terms the literature uses for your idea. That vocabulary is what makes your database searches work.
- Never let a reference reach your document straight from a chat. Every candidate goes through a database first. Resolve the DOI at doi.org, then search Crossref, your library discovery service or Google Scholar by exact title plus first author. We walk through the checks one by one in how to check whether a source really exists.
- Ignore the surface signals. In the reference study, links were more likely to appear inside fabricated citations than real ones, and about a third of the links attached to real works were inaccurate (Walters and Wilder, p. 5). A string that has the shape of a DOI is not a DOI until it resolves.
- Open the source and find the passage. A citation is a promise that a specific page says a specific thing. Take the page number from the printed page of the edition you read, not from your PDF viewer's counter; the page number guide covers scans, articles and ebooks with no fixed pagination.
- Format last, and check the whole list before you submit. Style errors are cosmetic and fixable in minutes, while a reference that turns out not to exist forces you to rewrite whatever argument rested on it. Run the finished bibliography through the bibliography checker so that every entry is confirmed against a database in one pass, and use find source for the opposite direction, when you have a paragraph and need a real work that supports it.
The academic literature cited in this article was pulled from cytado's own corpus, which is built from Crossref, OpenAlex, DOAJ and Biblioteka Nauki, with page numbers read off the source documents. We do not write reference lists from memory either.
What not to do
- Do not paste a bibliography you have not opened. If you have not seen the paper, you cannot claim it says anything.
- Do not treat plausibility as evidence. Fabricated references tend to resemble authentic papers with similar titles and overlapping author lists, which is exactly why they pass a glance (Chelli et al., p. 8).
- Do not use a general-purpose chatbot as your primary or exclusive search tool for a literature review; on the evidence of the systematic-review comparison, that is explicitly not recommended (Chelli et al., p. 1).
- Do not fix a suspect citation by asking the model for a corrected version. Look the work up in a database and copy the metadata from the record.
Frequently asked
- Is it cheating to use AI to find sources?
- Using a model to understand a debate, generate search terms or point you towards a literature is ordinary research practice, and most universities treat it that way as long as you disclose assistance where your institution requires it. What gets people into trouble is submitting references nobody verified. The rule that keeps you safe is simple: every source in your bibliography is one you located in a database and read.
- Can my supervisor tell that a reference came from a chatbot?
- Not from the writing style with any reliability, but a fabricated reference gives itself away the moment somebody searches for it, because it resolves to nothing. Wrong page numbers and wrong volume numbers on real papers are the other common tell, since those are the fields models get wrong most often. There is more on what supervisors actually catch in our article on whether your supervisor will spot AI-generated sources.
- I cannot find a paper my AI tool suggested. Does that prove it is fabricated?
- Not on its own, but it shifts the burden. Search the exact title in Crossref and Google Scholar, then search the first author's surname plus two distinctive words from the title, in case the title was mangled rather than invented. Check whether it might be a book chapter, a conference paper or a thesis rather than a journal article. If nothing surfaces after that, treat it as fabricated and drop it.
- Which AI model hallucinates the fewest citations?
- Reported rates vary enormously by model, by version and by field, from around 18% up to over 70% in published comparisons, and they change with every release. Picking a model on that basis still leaves you checking every reference by hand, so it is not a decision worth optimising. Verify the list, whichever model produced it.
Sources
Read next
- Will Your Supervisor Spot AI-Generated Sources?AI writing detectors are unreliable, but a fabricated reference either resolves in a database or it does not, which is exactly why fake sources are the part of AI-assisted work that gets caught.
- How to Check Whether a Source Really Exists (DOI, Crossref)Resolve the DOI, search Crossref by author and title, and read the reference for the tells that give away a fabricated citation.
- How to Spot AI-Fabricated Sources in Student WorkA practical guide for supervisors: how to recognize AI-hallucinated citations, verify them by hand, and why 30 theses demand batch verification.