AI is making science easier to access but harder to trust. Generative tools can now produce scientific-looking text, data, and citations at scale, yet verifying that content remains slow, expensive, and expertise-dependent. This growing gap between cheap generation and costly verification is creating new vulnerabilities in the scientific record and raising the stakes for scientific literacy.
The two fronts of the reliability problem
AI creates two distinct points of vulnerability. First, it can fabricate or manipulate research components-text, images, datasets, citations-that may later enter the published record. Second, it can distort genuine research when findings are selectively summarized or explained to readers. The first concerns the integrity of what gets published. The second concerns whether existing research is represented accurately.
Recent studies quantify both risks. A 2026 paper by Asai et al. found that GPT-4o, tested without external retrieval, fabricated citations in 78-90% of cases when asked to cite recent scientific literature. Yet OpenScholar, a specialized system that retrieves information from millions of open-access papers, achieved citation accuracy comparable to human experts. Another study by Topaz et al. (2026) examined 2.5 million biomedical papers in PubMed Central's Open Access collection and reported a rise in suspected fabricated references between 2023 and early 2026. Misleading citations are passing existing editorial checks and entering the record.
When the signals of credibility mislead
Readers and editorial systems often rely on familiar credibility signals: manuscript structure, methodological detail, institutional affiliations, polished figures, and consistency between sections. AI tools reproduce these features because they are trained on extensive collections of scientific text. But reproducing the patterns of scientific writing does not demonstrate that the underlying research occurred.
AI can also generate multiple research components that appear mutually supportive. A fabricated dataset may seem to back a fabricated figure, which in turn appears to support a fabricated conclusion. This internal consistency can make misleading material harder to detect. Reliable research requires authentic data, appropriate methods, transparent reporting, and identifiable people who remain accountable for the work.
Misrepresentation of genuine studies is another risk. An AI summary might cite an observational study linking supplement use to lower dementia rates but present that association as evidence of prevention. The reference is real, but the explanation turns a correlation into a causal claim. Checking that the citation exists would not catch the distortion. AI may also overlook methodological weaknesses or combine conflicting evidence into a false impression of consensus.
Detection is not verification
Detecting whether text was probably generated by AI is not the same as verifying whether the research is reliable. A manuscript may contain AI-assisted language yet report sound research. A text written entirely by a person may contain fabricated data or references that do not support its claims. AI detection is probabilistic and generates false positives. Verification focuses on whether the study took place, the data are authentic, and the conclusions follow from the results.
Editorial checks should therefore not be built primarily around identifying AI-generated text. They should strengthen verification of the research components that matter, regardless of how the text was produced. Publishers must also determine which checks are needed and when, with closer examination of references, images, or data where risks are higher.
The burden of checking
Expecting readers to validate content independently is unrealistic. Verification requires access to sources, methodological knowledge, time, and often specialist tools. As AI makes scientific-sounding content faster and cheaper to produce, the cost of verification does not fall at the same rate. This creates a growing burden for editors, reviewers, and researchers.
Review by subject-matter experts remains essential, but expertise is scarce and reviewer capacity is already limited. Stronger editorial checks can help prevent unreliable research from being published, but poorly designed or overly burdensome checks may also block legitimate work. Verification has an operational cost that requires investment in specialist staff, infrastructure, and editorial expertise as submission volumes and complexity increase.
For professionals working in AI for Science & Research, these dynamics are reshaping how institutions approach research integrity. The challenge is not simply detecting AI involvement but building verification workflows that keep pace with generation speed. Those developing or deploying AI tools in research settings increasingly need to understand both the capabilities and the failure modes of these systems, which is why structured AI Research Courses are becoming a practical requirement rather than an optional supplement.
Why this matters for science and research professionals
Scientific literacy now requires asking questions that go beyond the text itself: Where did this claim originate? What kind of evidence supports it? Has it been peer reviewed? What uncertainties remain? Has the research been corrected or retracted? Has AI transformed or selectively presented the original information? None of this requires every reader to become a methodological expert. It requires that the origin, status, and limitations of research be visible and understandable-including when findings are presented through AI. Publishers and AI providers share responsibility for making evidence easier to inspect, limitations easier to understand, and corrections harder to miss.
Your membership also unlocks: