News

This AI tool claims to pick the top 1% of preprints. Should researchers trust it?

Nature
10 August 2026

QED Science says that its metrics reduce bias by assessing papers solely on the basis of their originality and validity.  

A growing number of private firms are offering researchers artificial-intelligence tools for scrutinizing manuscripts before publication. One such system, developed by a start-up called QED Science in Tel Aviv, Israel, aims to judge whether life-sciences research is original and valid.

QED Science’s tool is trained to assess whether the claims in a given manuscript are supported by the data presented, and to identify gaps in the work. It is available to researchers for free, and has so far been used by more than 10,000 laboratories across 1,500 institutions in more than 70 countries.

In November last year, openRxiv — the non-profit organization that operates the preprint servers bioRxiv and medRxiv — announced it would be piloting the QED Science system on bioRxiv.

In an analysis posted in June, QED Science used the tool to rank more than 57,000 preprints that were posted on bioRxiv between May 2025 and April 2026, and selected the ‘top 1%’.

But the ranking has sparked debate among researchers, with some arguing that it risks creating another badge of prestige — and reinforcing metric-based culture in academia. Others have expressed doubts about AI's ability to reliably and transparently judge the quality of scientific research.

Nature spoke with Niv Mastboim, co-founder and chief executive of QED Science.

How does QED Science’s AI platform assess the science claims in a given article?

There are a lot of tools for reviewing research articles. We go about it a bit differently, in a couple of aspects. We have trained the AI platform on multiple data sources, from open reviews and user feedback to synthetic data.

One of the elements is establishing what would have been the negative results of the experiments — results that do not support the original hypothesis being tested, or the conclusion.

The published literature disproportionately represents successful and positive findings. To evaluate scientific claims properly, a system also needs to understand what evidence that fails to support a claim looks like. This can be learnt from published null or contradictory findings, failed replication studies and other forms of non-supportive evidence.

We focus on creating internal metrics for the system to constantly improve. So, the system validates each of the claims made in a paper and attaches a score to it. We use multiple scoring, from originality to validity.

The AI platform is completely autonomous, but we also have a lot of users who give us feedback. We use that feedback to tune the tool constantly.

What purpose did the top-1% list aim to serve?

The goal was to judge science on the basis of the work itself, not the journal or the prestige of the authors. The 574 preprints in the 1% are the top-scoring bioRxiv preprints among the 57,455 assessed. They were selected solely according to QED’s assessment of originality and validity, independently of author identity, institution or publication venue.

We also aimed to identify amazing papers that have been missed by the current publishing system. We did a separate validation analysis involving 2,879 bioRxiv preprints from April 2025 that were subsequently published in peer-reviewed journals, and then we benchmarked the ranking of our tool against journal ranking.

In that analysis, QED rated 12.9% of these papers more highly than their eventual journals of publication might suggest. We call these articles ‘hidden gems’. We consulted a panel of experts to judge, in a blinded manner, the strongest cases of disagreement, and found that they preferred the QED-favoured paper in 75% of decisive comparisons.

Is the tool the ultimate ‘peer reviewer’? Do you see it being used by publishers in the future?

We are not offering our products to journals or publishers. Our goal is to provide free services to authors in a private, secure environment so they can improve the work before it is published.

People want to improve their research, but they are lacking good critical judgement to help them do that. We are not going to replace their judgement; we’re going to augment it and help them spot things that would otherwise have been missed.

What about the other 99%? Does the platform suggest they should be ignored?

We completely acknowledge the notions that have been brought up by the research community — that we’re creating a scarcity because only a limited number of papers can reach the 1%. The goal was, let’s all agree this is amazing science.

That does not mean that there is not amazing science in the 98th or 97.5th percentiles. And in fact, we’re aiming to release multiple scores across multiple dimensions.

Do your rankings reinforce the culture of prestige metrics and publish-or-perish pressures in academia?

We do not want to create an arms race. The goal is to produce good science. Researchers can use the platform to see which claims are less or more fragile, and what information or explanation they should add. They now have a way to go about this before they release their findings to the world.

Read the full article
No items found.

Free access for academic researchers

Create your free QED account to validate your research, strengthen grant proposals, and uncover scientific insights.
We've sent you an access link.
Please check your inbox.

Didn't get your email? Check your spam folder or reach out to info@qedscience.com

Oops! Something went wrong while submitting the form.
Looking for QED for pharma, biotech or life science organizations?