Key Takeaways
- Critique tests the evidence: It asks whether the evidence supports each claim, rather than summarizing the paper or placing it within a field.
- Four domains carry most of the risk: Originality, reporting completeness, statistical consistency, and bias in how the result reached publication.
- Citation counts measure attention: The rhetorical function of a citation determines what the underlying evidence is actually worth.
- AI can handle the mechanical work: It can support critique of your own manuscript, but confidentiality rules limit its use on manuscripts assigned for peer review.
A paper can be correct in every sentence and still not be true in the way its abstract implies.
The scale of the problem is easy to state when you consider that over 1.5 million biomedicine and life science articles are published and indexed in PubMed every year. No researcher reads a meaningful fraction of that, and no reviewer can check it independently.
Evaluation capacity is the binding constraint. A review that confirms a paper is well written, well cited, and published somewhere respectable will still miss the failures that cost real time. These include conclusions that travel further than its data, controls that were never run, and results that hold in one cell line and nowhere else.
Scientists who do this well stop reading a manuscript as a single argument. They break it into separable claims, pair each one with the evidence behind it, and judge them one at a time.
To critique a research paper well, you need to test how the argument was built, how well the evidence supports it, and where the work needs to improve.
What Does It Mean to Critique a Research Paper?
Three different tasks get treated as interchangeable. A summary compresses a paper into its main points, a literature review synthesizes many papers to describe where a field currently stands, and a critique isolates one piece of work and tests whether its conclusions are earned by its data.
The distinction matters because science does not certify conclusions as final, but rather accumulates claims that have survived attempts to break them and revises them as new evidence arrives.
So a critique asks how strong a specific claim is in its specific context:
- Under which model system
- Which cohort
- Which analytic choices
- With what assumptions holding
That framing lowers the bar in one way and raises it in another. A critique does not demand a perfect paper, since no paper is perfect. It does demand that you locate the exact point where a conclusion outruns the evidence offered for it.
The output is a map. When you finish, you should know which parts of the work you would build on, which you would treat as provisional, and which you would set aside entirely.
What a Strong Critique Should Cover
Most of the risk in a manuscript clusters in four places. Working through them in a fixed order keeps a critique from becoming a list of impressions.
Originality and Conceptual Advance
Start with how far the finding moves the field. Separate a genuine conceptual advance from a methodological refinement of something already known, since both are often written up the same way.
Then ask what specifically is new, and whether that novelty supports the breadth of the conclusions drawn from it. A narrow, well-executed result presented as a general principle is one of the most common gaps in otherwise solid work.
Reporting Completeness
Incomplete reporting is a leading reason a result cannot be evaluated or reproduced, and it is the easiest problem to check. Every major study design has a reporting guideline specifying the minimum an author must disclose, collected by the EQUATOR Network. Pull the checklist that matches the design and read for what is missing.
Four guidelines cover most life science submissions:
Items left out entirely, rather than marked as not applicable, usually sit right on top of the fragile part of the experimental setup.
Statistical Consistency
Recompute before you interpret. Check reported p-values against their test statistics and degrees of freedom, and confirm that subgroup percentages resolve against their denominators. A GRIM test asks a narrower question: whether a reported mean is arithmetically possible for the stated sample size.
This catches more than most people expect. Across more than 250,000 p-values in eight psychology journals, half of papers using null hypothesis significance testing contained at least one p-value inconsistent with its test statistic and degrees of freedom, and one in eight contained a grossly inconsistent value that may have changed the statistical conclusion. The scope there is psychology rather than life science, and most of these are arithmetic drift across a long document rather than misconduct. They still need correcting.
P-hacking is a separate signal. A cluster of p-values sitting just under 0.05 is a prompt to look at analytic flexibility, such as post-hoc exclusions, covariates added late, outcomes selected after the results were known. Treat it as a signal to investigate further before reaching a verdict.
Sponsorship and Publication Bias
Two different distortions get filed under the same heading. Publication bias is a property of the corpus because null and negative results are less likely to be published, so the visible literature overstates effect sizes. It says nothing about the internal quality of the paper in front of you, and it does change how much weight a single positive finding deserves.
Sponsorship bias is a property of the study. Cochrane’s review of the evidence found that industry-sponsored drug and device trials are 27% more likely to report favorable efficacy results and favorable overall conclusions about 34% more often than studies with other funding sources.
Funding source is a reason to look harder at specific things. These include comparator selection and dosing, whether reported outcomes match the registered protocol, how adverse events are handled, and whether the conclusion is phrased more strongly than the results table supports.
A Practical Framework for Evaluating Scientific Claims
Evaluating scientific claims is a sequence, not a single judgment made at the end of a read. The five steps below move from the manuscript's internal structure outward to the literature it sits in.
- Decompose the manuscript into claims: Stop reading it as one narrative arc. List each assertion the paper makes and pair it with the specific evidence behind it, separating what was directly measured from what was inferred. Build sub-claims underneath each headline claim so the dependency structure becomes visible. This is claim tree extraction, and its main value is preventing persuasive writing from doing the work evidence should do.
- Audit internal consistency: Check degrees of freedom against sample sizes, percentages against denominators, and figure values against what the text says they are. Then check the logic: missing controls, assumptions the authors never stated, and alternative explanations that fit the same data equally well.
- Check reporting against the relevant guideline: Identify the study design, pull the matching checklist, and look for omitted items. An author who marks an item as not applicable has made a decision you can evaluate. An author who leaves it out has not.
- Test the claims against the outside literature: Ask whether the findings support, extend, or contradict what is already published. Then check whether the authors engaged with failed replications and null results in their own field, or cited only the work that agrees with them. Hold a claim to the standard of the evidence around it, not only the evidence inside the paper.
- Use AI for the mechanical load, within the confidentiality limit: Multi-agent systems can check figure-to-text consistency, recompute reported statistics, and surface contradictions with prior literature far faster than manual review. They can also produce an anonymized quality score for originality and validity, independent of journal placement or author affiliation. The limit matters because it applies to your own manuscript before you submit it. NIH prohibits peer reviewers from using generative AI to analyze grant applications or formulate critiques, and most major journals prohibit uploading assigned manuscripts to AI platforms on confidentiality grounds. Pre-submission self-critique and assigned peer review are different activities under different rules, a distinction most discussions of AI peer review skip.
researchers
Why Citation Context Changes How You Read a Claim
A citation count answers one question: how much attention a paper received. But it does not distinguish a citation that replicated a finding from one that failed to replicate it. Both increment the same number by one.
Automated classification tries to fix this by sorting citing sentences into supporting, mentioning, and contrasting categories, and the overwhelming majority land in mentioning. That distribution deserves some skepticism. In one sample of 98 citations drawn from systematic reviews, an automated citation index labeled 96 as mentioning and none as contrasting, while human coders assessed 42 as supporting, 39 as mentioning, and 17 as contrasting. The automated signal understates both agreement and disagreement, which matters most for the contrasting citations you would actually want to find.
The practical version is manual and takes about ten minutes. When a manuscript’s central premise rests on one heavily cited paper, go and read how that paper is actually cited. Pull a dozen citing sentences and read them in context.
If a meaningful share disputes the finding or reports failure to replicate it, the manuscript’s foundation is more fragile than its citation count suggests, and the authors may not know that. Prestige, citation count, and journal name are proxies. They correlate with quality well enough to be useful as a first filter, and they are not measurements of it.
How to Phrase Feedback Without Sounding Harsh or Vague
COPE’s ethical guidelines for peer reviewers require reviews to be objective, constructive, and respectful of the authors’ intellectual property. Vague dismissal fails that standard as clearly as a hostile tone does, and it fails the author too, because there is nothing in it to act on.
The mechanism is depersonalization. Point at the gap between a specific claim and specific evidence, name the standard you are applying, and state what would close the gap. Feedback structured that way is actionable whether or not the author agrees with it.
The difference shows up quickly in practice:
Each precise version names three things: the claim, the evidence, and the standard. That is the entire method.
Run the Framework on Your Own Manuscript First
Done by hand, this framework is slow, which is exactly why it gets skipped under deadline. Automating the mechanical passes, the arithmetic checks, the guideline cross-reference, the contradiction search, is what makes the judgment passes affordable.
Automated review has its own problems, and they are worth understanding before you trust one. Measuring whether a system’s comments are genuinely useful, and whether it catches everything it should, is an open engineering problem with real tradeoffs between precision and recall.
QED works on that problem directly, breaking manuscripts into individual claims and testing each against the literature rather than the sources an author happened to cite.
Run the framework on your own work before a reviewer runs something like it on your behalf. You innovate. We validate.
FAQs
researchers
researchers

