← All articlesWill Your Supervisor Spot AI-Generated Sources?
AI & sources2026-08-18· 6 min read

Will Your Supervisor Spot AI-Generated Sources?

AI writing detectors are unreliable, but a fabricated reference either resolves in a database or it does not, which is exactly why fake sources are the part of AI-assisted work that gets caught.

Probably, but not in the way most students fear. The software that claims to detect AI writing is unreliable enough that no serious supervisor should build a case on it. Your reference list is a different matter. Every citation is a factual claim about a document that either exists or does not, and checking one takes under a minute in a database that gives a yes or a no. That is why fabricated sources are the part of AI-assisted work that actually gets caught.

The AI detector is not the thing to worry about

The most comprehensive independent test of this software covered 12 public tools plus Turnitin and PlagiarismCheck, and concluded that the available detectors "are neither accurate nor reliable" and lean towards classifying output as human-written (Weber-Wulff et al., 2023, p. 1). The same team is blunt about what this means in a disciplinary hearing: because the tools produce no evidence, their reports "cannot be used as the only basis for reporting students for cheating" (p. 26).

The errors run in both directions. In a five-detector comparison, two genuinely human-written control texts were rated "Likely AI-Generated" by one tool, and the tools struggled far more with GPT-4 output than with GPT-3.5 (Elkhatat et al., 2023, p. 8). A Stanford group found that detectors "consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified" (Liang et al., 2023, p. 1). Weber-Wulff's team flags the same trap from another angle: running your own text through Google Translate or DeepL raises the false-positive rate, which puts second-language students at risk for something they did not do (p. 26).

Even the scoring is inconsistent across generators. When one plagiarism suite was fed introductions written by four different AI tools, the "likely AI-written" figure ranged from 27% for one tool to 100% for another on the same prompt (Lippi and Mattiuzzi, 2025). A number that swings that far tells your supervisor very little about you.

The reference list is checkable, and that changes everything

Verifying a citation is not a judgement call. Either a paper with that title, by that author, in that journal, in that year exists, or it does not.

The base rates explain why supervisors bother. Across 636 references in 84 AI-generated literature reviews, 55% of the GPT-3.5 citations and 18% of the GPT-4 citations were fabricated outright, and among the references that pointed to real works, 43% and 24% respectively contained substantive errors (Walters and Wilder, 2023, p. 1). Pooling six earlier studies, the same authors report 51% of 732 citations fabricated (p. 2). A comparative analysis of 471 references generated for systematic reviews found hallucination rates of 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Bard (Chelli et al., 2024, p. 1).

Newer models are better, not clean. A 2026 benchmark reports roughly 12% for GPT-4o, 16% for Claude 3.5 Sonnet and 21% for LLaMA-4 Maverick (Misra and Udandarao, 2026). At 12%, a bibliography of 25 items still contains about three references that will not survive a lookup.

What the check actually looks like

Supervisors who check do roughly what an automated verification pipeline does: exact lookup first, then a fuzzy search, then a judgement about whether the retrieved paper is really the one cited (Misra and Udandarao, 2026). In practice it is four steps.

Resolve the DOI. In a study of AI-generated musculoskeletal reviews, one prompt produced ten references of which four had DOIs that did not resolve at all and six were scored as fabricated (Safran and Çalı, 2025, p. 698). A dead DOI takes five seconds to find.

Search the exact title. Fabricated references "tend to look legitimate at first glance", which is precisely why they survive a skim and die in Crossref (Walters and Wilder, 2023, p. 2).

Read the numbers. Volume, issue, page range and year are where errors cluster; incorrect numeric values are the single most common defect even in citations to real papers (Walters and Wilder, 2023, p. 5). A quirk worth knowing: hyperlinks are more likely to appear in the fabricated citations than in the real ones, and about a third of the links attached to genuine works are wrong anyway (p. 5).

Open the paper and read the passage. This is the step that catches the harder cases. Emsley (2023), writing as a journal editor after testing the tool on his own field, found that the problem "goes beyond just creating false references" and includes falsely reporting the content of genuine publications. A real citation attached to a claim the paper never makes fails the same way a fake one does.

Before and after, from a bibliography that did not survive a check:

Before: Kowalska, M. (2019). Digital literacy and source evaluation among first-year students. Journal of Academic Skills, 14(3), 221–239. doi:10.1016/j.jas.2019.03.004

The DOI resolves to an unrelated paper, no journal by that name is indexed, and the author has published on a similar topic but never this title. That last detail is the signature pattern: fabricated content "often used real authors but invented titles or incorrectly matched publications" (Safran and Çalı, 2025, p. 700).

After: Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. https://doi.org/10.1038/s41598-023-41032-5

The DOI resolves, the record matches field for field, and the claim you hang on it is on a page you can name.

Asking the chatbot to check its own work does not help

The model that invented the reference will also confirm it. Walters and Wilder note that ChatGPT "often provides incorrect responses when asked 'Is this citation correct' or 'Do you fabricate citations?'" (2023, p. 2). Emsley describes the same behaviour as a tendency to double down convincingly when confronted. Verification has to happen against a bibliographic database, not against the generator.

This is also why the terminology matters. What looks like a hallucination is closer to a confabulation, an arbitrary and confidently wrong generation rather than a retrieval failure (Farquhar et al., 2024, p. 625); Emsley argues for calling the citation cases fabrications outright. Either way, the fix is external and mechanical.

The practical position

Using AI to draft, outline or rephrase is a conversation to have with your supervisor about their rules. Submitting references you have not opened is a different risk, and it is the one with a deterministic detection method behind it. Run your bibliography through a bibliography check before you submit, and if a claim in your text has no source yet, find a real one for that paragraph rather than asking a model to produce a plausible-looking line. If you want the manual version of the check, we wrote it up in how to check whether a source really exists.

Every academic reference in this article was pulled from cytado's own source corpus, with page numbers taken from the printed pages of the documents. That is the standard the tool exists to enforce, and the reason we can name a page for each figure above.

Frequently asked

Can my university prove I used AI to write my thesis?
Rarely, on stylistic grounds alone. The largest independent test of detection tools found them neither accurate nor reliable, and its authors state that detector reports cannot be the sole basis for reporting a student (Weber-Wulff et al., 2023, p. 26). Fabricated references are a different category of evidence, because a non-existent paper is a verifiable fact rather than a probability score.
My own writing was flagged as AI-generated. What now?
False positives are documented: human-written control texts have been rated 'Likely AI-Generated' by commercial tools (Elkhatat et al., 2023, p. 8), and detectors systematically misclassify non-native English writing (Liang et al., 2023, p. 1). Ask what the finding is based on, offer your drafts and notes, and point out that the tool produces no evidence beyond a number.
Newer models hallucinate less. Is it safe now?
Less is not none. A 2026 benchmark reports citation hallucination rates of about 12% for GPT-4o, 16% for Claude 3.5 Sonnet and 21% for LLaMA-4 Maverick (Misra and Udandarao, 2026). On a 25-item bibliography that is still roughly three references that will fail a lookup.
Is a reference safe if the DOI works?
Not automatically. A working DOI proves the document exists, not that it supports your sentence. Editors have documented AI output that cites genuine publications while misreporting what they say (Emsley, 2023). Open the page you cite and confirm the passage makes your claim.

Sources

  1. GPT detectors are biased against non-native English writers (Stanford HAI)
  2. Crossref metadata search
  3. DOI resolver

Read next