Quick answer: verify the claim, not just the reference
To verify AI research citations, split the draft into individually checkable claims, establish each source’s identity, open the relevant passage, and compare its scope with the exact sentence. Check versions and update notices before recording a support judgment and an editorial action. A real DOI establishes an identifiable work; it does not establish that the work supports your claim.
Keep three questions separate: Can we identify and inspect the source? Does its content support this wording? What should we publish? An inaccessible appendix is unresolved evidence, not a false finding. A corrected article can remain accessible while an older numerical claim is wrong. One source can directly support one sentence and only provide background for the next.
Start with a single short report. Use the 10-check worksheet and the complete fictional eight-claim source packet. The packet includes source passages, claim IDs, editorial decisions, and denominators. It is a teaching exercise, not customer research, an OpenMax product test, or a model-accuracy benchmark.
The review record: evidence status is not a support score
Treat an atomic claim as the smallest assertion that can be checked independently. “Adoption is high and the tool cuts costs” contains at least two. Separate them before assigning evidence. Keep the original sentence unchanged in the record; save the proposed revision in another field so the review does not silently erase its own starting point.
| Record layer | What to capture | What it does not prove |
|---|---|---|
| Claim identity | Claim ID, exact wording, draft version, surrounding paragraph | That a paragraph’s reference applies to every sentence |
| Source identity | Title, author or organization, publisher, DOI or stable URL | Accuracy, peer review, or support for the claim |
| Access and location | Full text, abstract only, or unavailable; section/table/page; short passage | That a search snippet is the full evidence |
| Current status | Source version, relevant date, correction/retraction notices checked | That absence of a notice proves no updates exist |
| Support judgment | Direct, partial, background, contradicted, or unresolved; reason | Overall source quality or permission to publish |
| Editorial action | Keep, rewrite, remove, or hold; owner; next check | That a suggested rewrite has already been approved |
Here, direct means the inspected passage supports the claim as worded, with its qualifications. Partial means a narrower part is supported. Background means topical relevance without evidence for the assertion. Contradicted means the inspected evidence conflicts with the wording. Unresolved means the evidence needed for a decision has not been obtained or interpreted adequately. These are this guide’s editorial labels, not a universal certification system.
A review entry should be reopenable without the original conversation. Record a stable section heading and table label as well as a page number: web pagination changes, and PDF viewer page numbers may differ from printed ones. Keep only the excerpt necessary for review and follow your organization’s permissions and retention rules. A saved link alone is not a reproducible evidence trail.
Ten checks that catch different citation failures
1. Establish that the cited work can be identified
Search the exact title, author, and relevant date. Resolve the supplied DOI and compare the destination with the reference. A journal name that exists does not prove a particular article exists. If the first search fails, try a shortened distinctive title, the publisher, a relevant scholarly index, and legitimate alternative versions. Record what was searched before calling the reference unresolved; a single empty search is not proof of fabrication.
A 2023 study by Walters and Wilder examined both fabricated citations and bibliographic errors in real works. That distinction is useful here; its historical model experiment is not a measurement of today’s systems. See the original research article.
2. Match bibliographic fields to the same object
Compare title, authors, venue, year, volume or article number, and identifier. Watch for references assembled from parts of two real papers. A preprint and a later journal article can have different identifiers and dates; do not combine the preprint’s findings with the journal version’s citation without inspecting the relevant version.
On September 4, 2026, we retrieved the public Crossref record for DOI 10.1038/s41598-023-41032-5 and compared the title, authors, journal, and publication date with the publisher page. They matched: Walters and Wilder, Scientific Reports, September 7, 2023. This was a bibliographic check, not a reproduction of the paper’s experiment. Crossref documents the metadata lookup; metadata is not full-text evidence and can itself need correction.
3. Open the material that contains the claimed evidence
A search result, abstract, press release, and full article offer different amounts of context. If your statement depends on Table 4, inspect Table 4 and its notes. An abstract may be enough for a narrowly worded claim it explicitly states, but not for a subgroup calculation hidden in an unavailable appendix.
If access stops at a paywall, use authorized library access, a legitimate author manuscript, or a request to the source owner. Record any version difference. Do not bypass access controls or treat an AI reconstruction as the missing document. Leave the claim on hold when the needed material is unavailable.
4. Pair an exact passage with one exact claim
Read the paragraph before and after the proposed passage. Ask which words in your sentence are actually supported. A source explaining what citation checking means does not establish that a tool verifies 90% of references. A source discussing potential benefits does not establish a measured result.
For medical journal manuscripts, ICMJE asks authors to ensure references support the associated statements and to verify references against original sources or appropriate bibliographic records. This is medical publishing guidance, not a rule imposed here on every business report. The underlying comparison is still useful: inspect the statement and its evidence together. ICMJE reference guidance.
5. Recheck the population, denominator, period, and units
Write down the numerator and denominator before repeating a percentage. “18 of 30 respondents” is not “18 of 120 invited organizations,” and neither estimates the whole market without further assumptions. Keep geography, observation dates, exclusions, and unit of analysis visible. People, accounts, sessions, documents, and claims are not interchangeable counting units.
When a number is derived, preserve the formula and source cells or sentences. Distinguish a relative change from a percentage-point change. If definitions differ between sources, do not average them into an apparently precise figure. Reconcile definitions first or explain why a combined number would mislead.
6. Separate observation, causation, and extrapolation
A before-and-after pilot can describe what changed. Without an adequate design, it does not by itself establish what caused the change. Look for controls, sample selection, task changes, missing observations, and the authors’ limitations before writing “caused,” “proves,” or “will.”
Prefer “the observed mean fell from 10 to 8 minutes during the pilot” over “AI reduced completion time by 20%” when the source cannot isolate the effect of AI. Do not extrapolate an eight-team exercise to all customers. A cautious rewrite is not cosmetic: it changes the assertion to one the evidence can actually support.
7. Preserve quotation meaning and translation boundaries
Check quotations character by character where practical, including negation, qualifiers, punctuation that changes meaning, and ellipses. Removing “not,” “may,” or a limiting condition can reverse a result. A polished paraphrase can also overstate certainty even when it uses none of the original wording.
For multilingual work, store the original passage beside the translation and label the translation as yours unless it is an official version. Review the meaning again in each language. Do not put a translated paraphrase in quotation marks as if it were the original author’s exact words. Ambiguous technical language belongs with a qualified reviewer, not an automatic confidence label.
8. Use the version appropriate to the claim’s date
“Currently supports” requires a current specification; “supported in June” requires the June version. Preserve both dates when a source changes. Do not replace an old citation with a new URL while leaving the old claim and approval untouched.
A release-note example in our packet has a fictional 100-page limit in version 1 and a 60-page limit in version 2. The old source can support a historical sentence, but not the present-tense specification. Version checks are also necessary for datasets, working papers, dashboards, and internal policies—not only journal articles.
9. Inspect corrections, retractions, and expressions of concern
Check the publisher page and relevant update notices before relying on a source. For participating publications, Crossmark can expose updates; Crossref explicitly cautions that its presence is not a guarantee. No Crossmark badge or empty metadata update field should be interpreted as a clean bill of health. Crossmark documentation.
A correction, a retraction, and an expression of concern are different signals. Read the notice and determine which claim it affects. ICMJE’s medical publishing guidance recommends retaining clearly marked retracted articles with linked notices: an article can therefore remain reachable after retraction. Check status, not just HTTP success. ICMJE retraction guidance.
10. Assign an action and preserve the unresolved queue
“Checked” is too vague. Keep a supported sentence, rewrite a partially supported one, remove an unsupported assertion when it is unnecessary, or hold it while obtaining missing evidence. A contradicted claim may be repairable, but the new sentence must be checked again. Do not simply relabel the old version as direct.
Record reviewer, date, reason, revised wording, and release decision. If a citation supports several passages, an update should reopen all affected claim IDs. Track unresolved items separately from completed edits; otherwise a smaller final denominator can conceal difficult cases that were silently dropped.
A five-step workflow for one report
- Freeze the draft and inventory the claims. Save a version, assign claim IDs, and distinguish claims needing external evidence from clearly labeled editorial advice. Start with conclusions, numbers, quotes, and recommendations that affect decisions. Do not assume a sample review validates unreviewed claims.
- Retrieve sources through authorized routes. Match identifiers, open relevant text, and capture location, version, access level, and update-check date. Return an unresolved entry when retrieval fails. Do not ask the same model to invent a substitute reference.
- Review support against the original wording. Apply the ten checks. Let an assistant propose passages or likely mismatches, but require a person to inspect the passage and its context. Resolve disagreements with a second reviewer when the interpretation matters; store both reasons until decided.
- Edit and recheck changed claims. Keep original and revised sentences side by side. Recalculate numbers, restore qualifiers, replace obsolete specifications, and obtain missing passages. Use a separate release decision so an evidence judgment cannot silently become publication permission.
- Close the report with visible exceptions. Reconcile original claims, revisions, removals, and holds. Name the owner for remaining work and a review trigger, such as a source correction or changed product version. Reopen dependent claims when that trigger occurs.
For a one-off memo, a browser and the worksheet are enough. Recurring reports need a shared location for records and a clear handoff. For research synthesis beyond citation checking, use the AI research report template. If the evidence is trapped in a scan or difficult table, resolve the PDF extraction workflow first, then return to support checking.
Eight fictional claims: what changes after review
Every source, organization, release note, and numerical scenario in this exercise is fictional. Source IDs identify passages in the downloadable packet, not real publications or made-up DOIs. The exercise demonstrates editorial reasoning under supplied facts; it does not prove the quality of an external source or the accuracy of an AI system.
| Claim | Original assertion or problem | Support as drafted | Editorial action |
|---|---|---|---|
| C1 | 60% of all 120 invited organizations use the feature | Partial | Rewrite for the 30 respondents |
| C2 | AI caused a 20% time reduction | Partial | Describe the observed change without attributing cause |
| C3 | The current limit is 100 pages | Contradicted | Use version 2’s 60-page limit |
| C4 | A glossary proves 90% verification success | Background | Remove the numerical assertion |
| C5 | A result depends on an unavailable appendix | Unresolved | Hold until the required passage is accessible |
| C6 | The corrected result is 40/200, or 20% | Contradicted | Use corrected denominator 160 and 25% |
| C7 | A quotation removes “not” from a causal limitation | Contradicted | Restore the complete limitation |
| C8 | 30 of 120 invited organizations responded: 25% | Direct | Keep the response-rate wording |
C1 — A correct percentage attached to the wrong population. Source S1 says 120 organizations were invited, 30 responded, and 18 respondents used the feature. The calculation 18/30 = 60% is correct; the generalization to all invitees is not. Revised: “Of the 30 respondents, 18 reported using the feature (60%).” Add that overall prevalence among all invitees is unknown. The response rate, 30/120 = 25%, is a different measure.
C2 — A measured change without an isolated cause. S2 describes eight teams whose mean completion time moved from 10 to 8 minutes, with no control group and changed task mix. The relative decrease is (10−8)/10 = 20%. Keep the before-and-after description and the design limitation. The source does not establish that AI alone caused the decrease or that another team should expect it.
C3 — A genuine source, wrong version. S3-v1 dated June 1 permits 100 pages; S3-v2 effective September 1 replaces it with 60. For a September 4 present-tense claim, cite version 2 and say 60. A historical sentence about June can still cite version 1. Do not delete the earlier document: its date explains why the draft once looked plausible.
C4 — Relevance mistaken for measurement. S4 defines citation verification and contains no evaluation data. It can support a definition, not “90% of citations are verified successfully.” Remove that statistic from the report. Finding another generic authority on the subject would not repair the evidentiary gap; a measurable performance claim needs an appropriately designed and documented evaluation.
C5 — Missing access is an open question. S5’s metadata identifies an appendix, but its contents are not supplied. The draft attributes a 12-point subgroup improvement to Table A2. We cannot verify the number, subgroup, or meaning of “points.” Hold the sentence and request an authorized copy. Do not infer that the study is false or downgrade it to background just because access is missing.
C6 — A correction changes the denominator. S6’s original note used 40/200 = 20%. Its correction retains the numerator but changes eligible observations to 160. The current proportion is 40/160 = 25%. Record the correction and revised result together. This example is a correction, not a retraction, and should not be described as one.
C7 — One missing word reverses the conclusion. S7’s complete fictional sentence is “Results do not establish causation.” The draft quotes “Results establish causation.” Restore the source wording or write an accurate paraphrase. A spelling check would not solve this; a comparison with the original passage does.
C8 — A supported claim still needs its limits. S1 directly supports “30 of 120 invited organizations responded, a 25% response rate.” Keep it. Do not expand the sentence to describe adoption or the whole market. A source’s ability to support this claim does not automatically make C1 acceptable.
Count the original eight claims before editing: one direct, two partial, one background, three contradicted, and one unresolved. Seven have the relevant body text available, so access coverage is 7/8 = 87.5%. Direct support is 1/8 = 12.5% across all original claims, or approximately 14.29% among the seven assessable claims. State which denominator you use.
The action count is five rewrites, one keep, one removal, and one hold. The six retained or rewritten candidates still need a release decision. They are not evidence that the original draft was 75% accurate. These arithmetic checks describe this deliberately constructed packet only; they are not a target rate, customer result, or current model benchmark.
Choose an approach by review volume and consequence
Manual review fits a short report or sensitive interpretation. A reviewer opens original sources and fills the worksheet. It offers direct context but can miss repeated dependencies. Use stable claim IDs and a second reader for consequential disputes; stop when the needed expertise is absent.
Native reference-manager features help maintain bibliographies, identifiers, and formatting where the selected tool supports them. Confirm the actual features before relying on them. A correctly formatted bibliography still requires claim-level review. Do not buy a formatting tool expecting it to establish causation or source quality.
Scripted metadata checks can detect missing fields, DOI mismatches, duplicate records, and stale links at larger volumes. Keep requests within the provider’s current access rules and record failures separately. Deterministic matching can flag inconsistencies; it cannot infer that a table supports a sentence simply because their titles share words.
Agent-assisted review can propose atomic claims, retrieve permitted material, and suggest support labels in a configured workflow. Its output is a candidate assessment. Require visible passages, uncertainty, and human review; a second fluent answer without new evidence is not independent verification. Start with the packet before granting access to real documents.
Recurring operation adds assigned reviewers, versioned records, exception queues, and update triggers. It is useful when handoffs become the bottleneck, not merely because there are many URLs. If nobody can own an unresolved claim, adding automation will not create accountability.
Where OpenMax fits—and what this guide does not verify
OpenMax describes itself as a human–agent collaboration platform. That makes this a relevant workflow to evaluate when research tasks move between people and assistants. It does not establish that OpenMax has a native DOI checker, scholarly database connector, retraction monitor, citation-support classifier, or publication approval system. Those capabilities were not verified for this guide.
For a proposed pilot, prepare the frozen report, permitted sources, claim worksheet, grading examples, and a named reviewer. Ask the team to demonstrate whether the configured environment can retain the original wording, show exact passages, distinguish unavailable evidence from contradictions, preserve edits, and return unresolved work to a person. These are acceptance requirements, not statements that the product already performs them.
Keep a separate source of record until those requirements are demonstrated. Do not give a pilot permission to publish or alter a live knowledge base. A one-off report with a capable reviewer may not need a collaboration platform at all. The smallest next step is to open OpenMax with the eight-claim packet and evaluate the handoff, rather than assuming that a generated reference list is a finished review.
Limitations and release rules
This guide is produced for OpenMax and discusses its potential fit; it is not an independent product evaluation. The fictional exercise and narrow bibliographic lookup are disclosed separately. No customer performance, expert endorsement, comprehensive retraction search, or independent reproduction is claimed.
Direct support is not the same as a strong study. A weak or biased source can accurately be quoted and still be unsuitable for a major decision. Evaluate method, relevance, conflicts, and corroboration in proportion to the claim’s consequence. Several articles repeating the same press release are not several independent measurements.
Do not upload confidential manuscripts, licensed full texts, personal information, or internal reports to unapproved services. Use authorized access and a minimum necessary excerpt; obtain appropriate privacy, security, or legal review for your circumstances. Medical, legal, financial, and other high-stakes conclusions require qualified subject-matter review. This editorial checklist does not supply that approval.
If a source changes after publication, identify affected claim IDs, suspend or qualify the affected statement where appropriate, obtain a fresh review, and record the correction. Preserve the older review so the change is intelligible. “Last checked” is a record of an observation, not a promise that a page will remain accurate indefinitely.
Frequently asked questions
Does a working DOI prove an AI citation is correct?
No. It helps identify a work. Match its metadata to the reference, then inspect the exact passage and relevant version. The work may exist while the citation names the wrong authors or supports a different claim.
Can an abstract be enough to verify a claim?
Sometimes, for a narrow assertion explicitly contained in the abstract. Record abstract-only access and do not extend the judgment to tables, subgroup results, methods, or limitations you have not inspected. Hold claims that depend on unavailable details.
Can another AI model verify the first model’s references?
It can help identify candidate issues, but agreement between models is not independent source evidence. Require the original passage, locator, and reason. A person must resolve consequential ambiguity and approve the wording before publication.
What should we do if a source is corrected or retracted?
Read the notice, identify the claims it affects, and reopen their review. A correction may require a numerical or wording change; a retraction can make reliance inappropriate except when discussing the retraction itself. Seek domain review rather than treating every notice as the same event.
How should we report the citation-verification rate?
Define the unit and denominator first. Report access coverage, original claim-support judgments, and editorial actions separately. In this fictional packet, seven of eight claims have relevant body text, but only one original claim is directly supported. Edited candidates must not be counted as originally correct.
Sources and next review
Primary sources were checked on September 4, 2026. The 2023 study is historical evidence about its own experiment; the two ICMJE links concern medical publishing. None of these organizations reviewed or endorsed the fictional exercise, OpenMax, or this guide.
- Crossref REST API documentation: bibliographic metadata lookup and its role.
- Crossmark: publication updates and the limits of the indicator.
- Walters and Wilder’s original research: fabricated references versus errors in real references.
- ICMJE manuscript reference guidance: checking references in medical manuscripts.
- ICMJE retraction guidance: notices and continued availability of marked articles.
- OpenMax homepage: first-party positioning, not verification of the proposed workflow.
Recheck this workflow when source access, publisher update behavior, or the proposed product configuration changes. Begin with the worksheet, compare decisions using the source packet, and bring unresolved interpretations to the person responsible for the report.

