Research integrity is the adherence to professional standards of honesty, accuracy, and accountability across the full research lifecycle, from study design and data collection through analysis, authorship, and publication. It is broader than research ethics, which governs the treatment of human and animal subjects. Integrity governs the truthfulness of the record itself.
The scope covers accurate data handling, transparent methods reporting, appropriate attribution and authorship, disclosure of conflicts of interest, responsible peer review, and correction of the record when errors surface. Failures here propagate: a fabricated dataset does not stay in one paper, it enters meta-analyses, guidelines, and downstream research programs.
Integrity is usually enforced against the well-known triad of fabrication, falsification, and plagiarism, but institutional codes reach further. Data management and retention, pre-registration fidelity, honest reporting of null results, reviewer confidentiality, and refusal to inflate authorship all fall inside the boundary. The practical test is whether an independent reader, given the same materials, would reach the same conclusions the authors reported.
Fabrication and falsification. Fabrication invents data outright. Falsification manipulates real data or images by trimming outliers without justification, duplicating gel bands, or adjusting figures selectively. Image manipulation is now the single most frequently detected form of research misconduct in the life sciences.
Plagiarism and text recycling. Beyond copied prose, this includes reusing one's own published text without disclosure and appropriating ideas encountered during confidential peer review.
Authorship abuse. Gift authorship, ghost authorship, and the omission of contributors who did the work. Authorship disputes are the most common integrity complaint institutions receive.
Undisclosed conflicts and paper mills. Failure to declare funding or commercial interests undermines interpretation. Paper mills industrialize the problem, selling authorship slots on fabricated manuscripts at scale.
Peer review was never designed as an audit. Reviewers assess plausibility and contribution using the manuscript alone, usually without raw data, code, or protocols, and they work unpaid under time pressure. A competently fabricated dataset looks entirely ordinary in that setting.
Detection is also structurally delayed. Statistical impossibilities and image duplications typically surface years after publication, often through post-publication scrutiny rather than journal processes. Meanwhile incentives push the wrong way: publication volume drives hiring, promotion, and funding, while replication and correction carry no comparable reward. Add the questionable-practice grey zone, where selective outcome reporting and post-hoc hypothesis framing are common and rarely investigated, and most of the problem sits below the threshold anyone formally reviews.
Automated screening now runs at a scale human reviewers cannot match. Image forensics tools detect duplicated and spliced panels across entire journal archives. Statistical consistency checks recompute reported test statistics against degrees of freedom and flag impossible means for the given sample sizes. Text models identify paraphrased boilerplate and the distorted synonyms characteristic of paper mill output.
The more consequential shift is from surface screening to claim verification: systems that trace each stated conclusion back to the evidence presented and evaluate whether the inference actually holds. That catches a different failure class, including conclusions unsupported by the reported effect and citations that do not say what the citing paper claims. QED Science's work on AI infrastructure for scientific validation targets this layer.
The limits are real. Automated flags are signals, not findings, and false positives carry serious professional consequences. Every credible workflow routes flagged items to human adjudication before any allegation is made. Applying a validated quality metric at the pre-submission stage prevents more damage than detection after publication does.
Pre-register and version the protocol. A time-stamped analysis plan removes the ambiguity that makes selective reporting invisible.
Share data and code by default. Deposited datasets with persistent identifiers make claims checkable and deter fabrication more effectively than any policy statement.
Formalize authorship early. Use CRediT contributor taxonomy and agree on order at project start, not at submission.
Build internal pre-submission review. An independent check on statistics, figures, and claim-evidence alignment inside the lab catches errors while they are still correctable.
Fund the infrastructure. Institutions need a trained research integrity officer, a protected reporting route, electronic lab notebooks with audit trails, and consequences that apply to senior faculty as well as trainees.