← All articlesCan Professors Tell Your Sources Came From AI? (And How They Catch It)
AI & sources2026-07-31· 5 min read

Can Professors Tell Your Sources Came From AI? (And How They Catch It)

Markers do not detect an AI writing style — they check whether your references exist. Here are the measured hallucination rates and the five signals that give a fabricated bibliography away.

Yes, they can — but not in the way most students expect. Your professor is not detecting an "AI voice" in your prose, and does not need a text detector to do it. They take three entries from your reference list and try to find them. The whole advantage sits with the person checking, because of one asymmetry: proving that a reference exists is trivial, while proving that one does not exist is time-consuming at best and impossible at worst (Glynn, 2025, p. 3). Which means the burden falls on you. You produce the source; nobody has to prove its absence.

Below: how often models actually fabricate citations, the specific signals markers act on, and what to do before you submit.

The scale: this is not the occasional slip

Citation hallucination is measurable and common. In a study comparing model-generated literature searches against human-conducted systematic reviews, the proportion of fabricated references was 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Bard (Chelli et al., 2024, p. 5). The criterion was strict but reasonable: a paper counted as hallucinated if at least two of three fields — title, first author, or year of publication — were wrong (Chelli et al., 2024, p. 1).

The more revealing finding concerns the references that were not fabricated. Even among those genuine papers, a correct DOI was supplied in only 16% of cases for GPT-3.5 and 20% for GPT-4 (Chelli et al., 2024, p. 5) — against near-perfect accuracy on the titles themselves. So a model can correctly name a real article and attach an invented identifier to it. That is exactly why an AI-assisted bibliography collapses on the first click of a DOI, even when the titles look entirely credible.

A second study, on AI-generated text about musculoskeletal rehabilitation, was harsher still. Of 40 references analysed in the initial generation, only 7.5% were fully accurate, 42.5% were completely fabricated, and the remaining 50% were partially correct (Safran and Çalı, 2025, p. 695). The most common defects were invalid DOIs, fabricated article titles, and mismatched metadata (Safran and Çalı, 2025, p. 695).

Rates vary by model and by field — reviews report figures from roughly 18% for GPT-4 to over 70% for other frontier models, reaching 88% in legal contexts (Misra and Udandarao, 2026). None of those levels is survivable in a dissertation.

What actually gives it away

The signals markers act on are concrete:

  1. A DOI that will not resolve. Paste it into doi.org — that is the entire test. Given that a correct DOI appeared in only about one in five genuine entries, this alone breaks most AI-assisted reference lists in seconds.
  2. An entry that appears nowhere. Not in Google Scholar, not in Crossref, not in the library catalogue. One such entry is an accident; three is a pattern.
  3. An author-journal mismatch. A recognisable name attached to a periodical they have never published in — and your supervisor typically knows both the names and the journals in their own field.
  4. Traces of the chat window. Published bibliographies have repeatedly been found containing a pasted ChatGPT interface label; in at least 50 instances, authors carried the phrase "Regenerate response" into their manuscripts (Glynn, 2025, p. 3).
  5. A request for the PDFs. The most effective move available: "send me the full text of these five." You cannot send what does not exist. This is precisely the logic behind the proposal that journals require authors to deposit the full text of every cited source alongside the manuscript — cheap for editors, and effectively immunising against hallucinated references (Glynn, 2025, p. 2).

What happens when it surfaces

It is worth seeing how this plays out beyond the classroom. In March 2024, PLoS One published an article whose reference list was challenged by commenters on PubPeer. The paper was retracted 45 days after publication: the editors reported that 18 of its 76 references could not be identified, and treated the stray chatbot interface phrase as evidence of potential undisclosed AI use (Glynn, 2025, p. 3).

For a term paper or dissertation the consequences vary — a resubmission, a capped mark, or an academic misconduct hearing — but the mechanism is identical. What gets questioned is the credibility of the whole submission, not one entry. If three references turn out not to exist, nobody has reason to trust the remaining thirty.

What to do before you submit

The good news is that this is fixable, and quickly. In the Safran and Çalı study, structured post-generation verification raised the proportion of fully accurate references from 7.5% to 77.5% (2025, p. 695). The checking process — not a better model, not a cleverer prompt — accounts for that entire gain.

In practice that means three habits:

  • Resolve every DOI. If it lands on a different article than the one in your reference list, treat the entry as unverified.
  • Get the file. The minimum standard: nothing stays in your bibliography unless you have the PDF or library access to it.
  • Check the list in bulk. By hand it is two to four minutes per entry. Our bibliography checker queries identifier databases and shows which entries are confirmed, which have mismatched metadata, and which cannot be found at all. The underlying procedure is covered in how to check if a source really exists.

When an entry does fail, you need a replacement that supports the same claim — that is what the source finder is for, returning real publications with citable page numbers.

Every study cited above was retrieved by cytado itself, from its own source corpus, with page numbers taken from the original files' pagination. Writing about fabricated citations using fabricated citations would be a poor look.

Frequently asked

Do professors use AI text detectors?
Sometimes, but that is not where the risk lies. Style detectors are unreliable and contestable, whereas checking whether a cited publication exists is decisive and takes seconds. A reference list is simply far easier to verify than prose.
How often do language models fabricate citations?
In a comparative study the hallucination rate was 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Bard. Broader reviews report roughly 18% to over 70% depending on the model, and as high as 88% in legal contexts.
The title is right but the DOI does not work. Is that a problem?
Yes, and it is the typical failure. In the same study, among references that were not fabricated, a correct DOI was given in only 16-20% of cases despite near-perfect title accuracy. Models routinely name a real article and attach an invented identifier. Find the real DOI and correct the entry.
Can I fix this before submission?
Yes. Structured verification after generation raised fully accurate references from 7.5% to 77.5% in one study. The gain comes from the checking process itself, not from switching models or writing a better prompt.

Read next