Key Takeaways
- Opacity has three main drivers: Unpublished reviewer reasoning, documented prestige bias, and undisclosed AI use. In one recent global survey, more than 50% of reviewers reported using AI while reviewing manuscripts.
- AI and experts play different roles: AI provides structured, evidence-linked first-pass coverage. Experts provide judgment, context, and accountability. Transparency comes from recording how the two interact.
- Policies are converging: Major journals, funders, and conferences increasingly allow AI assistance within stated limits, while keeping judgment human and requiring disclosure.
In 2025, Nature, the world’s most cited scientific journal, decided its own review process had been too hidden to defend.
Researchers know all too well how evaluation decisions shape careers, funding, and entire research directions. Yet the reasoning behind most of those decisions never becomes visible to the people they affect, let alone to the readers who build on the published result.
A second layer of opacity is now forming on top of the first. Reviewers are quietly using AI tools that reports don't disclose and that policies are still catching up to. We cover both in more detail in the next section. But research evaluation has two visibility problems at once: hidden human judgment and hidden machine assistance.
The response forming across journals, funders, and conferences is a move toward evaluation that leaves a record. The interesting question is what belongs in that record, and what AI and human experts each contribute to it.
The strongest approach combines what each side does well into a workflow that a third party can inspect, with real-world examples showing where it is already in use.
Why Research Evaluation Has a Transparency Problem
For most of its history, peer review has run on a confidentiality bargain where reviewers speak freely because no one outside the process will read what they wrote. The cost of that bargain is that the reasoning behind acceptance, rejection, and funding decisions disappears.
In June 2025, Nature made transparent peer review standard for all new submissions, publishing reviewer reports and author responses alongside accepted papers. When a journal of that standing changes a longstanding practice, it is acknowledging that peer review transparency had become a credibility problem. Naturally, the rest of the field noticed.
Hidden reasoning also hides bias. In a controlled experiment at a computer science conference, reviewers who could see author identities were significantly more likely to recommend papers from famous authors and high-prestige institutions than reviewers scoring the same papers blind. No published record showed those decisions were influenced. Without a record, bias is undetectable, and what cannot be detected cannot be fixed.
The newest driver is machine assistance nobody declares. In a survey of roughly 1,600 academics across 111 countries, more than 50% of reviewers reported using AI tools while reviewing manuscripts, often against the guidance of the venue they were reviewing for.
Machine learning conferences are already policing this. ICML 2026 desk-rejected just under 500 papers, about 2% of submissions, after catching reviewers who had agreed to a no-LLM policy and used one anyway. Concealment drew the penalty, while reviewers who declared limited AI use under the conference’s alternate policy faced none.
What AI Is Good At in Research Evaluation
The case for AI peer review assistance rests on properties that hold at the fortieth manuscript of the week as well as the first.
When Stanford researchers compared GPT-4 feedback with human reviewer comments across roughly 3,000 papers from Nature-family journals, the overlap between model and human was comparable to the overlap between two human reviewers. The model raised many of the same substantive points the experts did.
Still, these four properties do most of the work:
- Coverage at scale: A model reads every section and checks every claim with the same attention, whether the manuscript is the first of the day or the fortieth. Reviewer fatigue has no machine equivalent.
- Consistency: The same criteria get applied to every submission. Two manuscripts making similar arguments get examined the same way, which is a precondition for fair comparison.
- Structured output: Findings arrive mapped to specific claims and the evidence behind them. This is the property that matters most for transparency, because a structured finding can be inspected, disputed, and overruled, while a vague impression resists all three.
- Speed: A first pass arrives in minutes. When journal decisions take months, feedback that arrives while the author can still act on it changes what feedback is for.
researchers
What Expert Human Review Is Good At
None of the above adds up to judgment. Deciding whether a finding matters, whether a methodological shortcut was reasonable given the constraints of the field, or whether a claim is important enough to justify the evidence offered for it requires context that lives in a reviewer’s head, accumulated over years of doing the work.
The published evidence supports this division. The same analysis found the model leaned toward generic suggestions and struggled to critique method design deeply. Surface-level gaps get caught, but the flaw that actually sinks a paper can slip through.
Constraint awareness fits into the same picture. A grant reviewer knows what a funding call is actually asking for and what a page limit forced the applicant to cut. Language models can penalize work for omissions the format required, a failure mode any researcher who has run a proposal through one will recognize.
Researchers themselves have been consistent about which parts of review they want kept in human hands. In Wiley's ExplanAItions 2025 report, drawn from surveys of roughly 5,000 researchers, nearly all peer-review use cases landed in the humans-preferred quadrant, with about 60% saying humans currently outperform AI on them. Funders have drawn the line harder, with NIH prohibiting its reviewers from using generative AI to analyze applications or draft critiques, citing confidentiality and integrity. A judgment someone signs and can be questioned about carries a weight no system output has on its own.
How the Two Combine: A Transparent Evaluation Workflow
The workflow that makes research evaluation inspectable puts these strengths in sequence, with a human in the loop at the decision point. The AI produces a structured, evidence-linked first pass.
The expert adjudicates it, accepting, rejecting, or contextualizing each finding. Every step leaves a record.
The quality of the first pass is measurable, and teams building these systems validate it as any instrument is validated: by measuring the precision and recall of AI-generated review findings against expert judgment.
Publication ethics bodies supply the final requirement: COPE’s position is that AI cannot be an accountable author and that its use must be disclosed, a principle that extends naturally to review. The human adjudicator owns the outcome.
Here is what each step contributes to the record:
Where This Is Already Happening
Pieces of this workflow are already running at scale, in venues that publish their results. These include:
- Journals: Nature Portfolio’s transparent peer review rollout puts reviewer reports and author responses into the public record across a growing set of journals. That is the disclosure half of the workflow, operating at the top of the prestige hierarchy.
- Conferences: At ICLR 2025, a machine learning conference, a randomized controlled study across more than 20,000 reviews delivered LLM feedback on draft reviews. 27% of reviewers updated their reviews in response, and blinded evaluation found the revised reviews more informative. The human reviewers remained the decision-makers throughout.
- Publishers and funders: Policies from major publishers now converge on three commitments: assistance permitted within stated limits, judgment human, use disclosed. NIH’s reviewer prohibition sits at the strict end of the same spectrum.
Put a Structured First Pass in Front of Expert Judgment
This division of labor is the design principle behind QED Science’s review products.
Paper Review and Grant Review break a manuscript or proposal into its individual claims, evaluate the evidence behind each one after full anonymization, and return structured findings for the researcher to adjudicate.
The same approach powers QED Score, a validated AI-based quality metric for life-science manuscripts, built to strengthen expert judgment and leave every finding open to inspection. That signal follows you into the daily reading problem too, where the Insights Feed surfaces new literature relevant to your own work, with QED Score data alongside its top-scoring recommendations.
Upload a manuscript or grant before your next submission and see what a claim-level first pass surfaces while there is still time to act on it. You innovate. We validate.
FAQs
researchers
researchers

