When a colleague asks me what I think of a paper, I've noticed I almost never answer with a verdict on the paper itself. I answer with a verdict on a claim inside it. "The knockdown data are solid, but I'm not convinced the phenotype is caused by the mechanism they propose." Two different judgments, bundled into one document, and only one of them is what the discussion is actually about.
That habit, I think, points to something the way we talk about scientific quality tends to obscure: the paper is not the unit of science. The claim is.
Knowledge, prediction, product
Science exists to advance what we know about the world, what it's made of, how it came to be, how it operates. We generate that knowledge through hypothesis and experiment, and the output of a well-run experiment, if it succeeds, is a claim. "Gene X is responsible for trait Y" is a claim. It is small, falsifiable, and crucially it is the thing that lets us predict something we haven't yet observed: that a drug activating gene X will affect trait Y.
This is the chain that gives knowledge its value. Knowledge lets us predict. Predictions, acted on, become products, whether that’s drugs or bridges, satellites or bolts. A product only works because the prediction behind it was accurate, and the prediction was only possible because the underlying knowledge was, to a sufficient degree, true. Technology, in this sense, is knowledge that has been turned into a working prediction.
Which raises the obvious question: how do we know, at the time a claim is made, how true it really is?
The trouble with truth
The thing is that usually, we don't, not directly anyway, and not any time soon. Truthfulness tends to be visible only in hindsight, through the predictions a claim goes on to enable and, often, through the better claims that eventually replace it.
Dalton claimed atoms were solid, indivisible spheres, something like billiard balls. That claim correctly predicted that elements combine in fixed weight ratios, a real and useful prediction. Decades later, Schrödinger's quantum-mechanical account of the atom explained and predicted chemical behavior that Dalton's model never could. Both claims were legitimate in their moment. Schrödinger's was simply closer to what's actually there.
Or take Hershko and Ciechanover's discovery of the ubiquitin-proteasome system, the mechanism cells use to actively tag and degrade proteins. When it was first described, its significance wasn't obvious. It took years before that mechanism became a validated drug target, now used in treating multiple myeloma. The claim's truthfulness only became visible once it had produced something that worked in the clinic.
This is the uncomfortable feature of truth as a criterion: we frequently cannot use it in real time, because we don't have it in real time. We usually find out later, sometimes much later, whether a claim was pointing at something real.
What we actually have: validity
So we substitute a proxy, and I think it's worth being explicit that it is a proxy: validity. Validity is not "is this claim true?" It's "how well is this claim supported by the evidence, given the tools and knowledge we have today?" It is dated, local, and testable in the present, exactly what truth is not.
The two move together, most of the time, but they are not the same thing, and the gap between them is where many of the interesting problems sit.
Validity can be high while truthfulness turns out to be limited. Newton's laws are about as well-supported as a claim in physics gets (we still use them to put satellites into orbit!) and yet quantum mechanics tells us they don't describe reality at small scales. Newton's laws were, and remain, valid. They are not fully true.
Validity can also decay without the claim itself changing at all. I can run the best experiment available today and validly support a claim. New instruments, better controls, a larger sample next year, and the same claim, unchanged, might fail to hold up under re-examination. Validity isn't a permanent verdict handed down once; it's a snapshot, and the snapshot has an expiration date tied to the state of the field.
This is why I've come to think validity, not truth, is the honest target for any evaluation of ongoing science. We are not, and cannot be, in the business of certifying truth in real time. We can be in the business of rigorously assessing how well a claim is currently supported, which is our best available estimate of how close it is to the truth, without waiting the years or decades it would take to know for certain.
No claim stands alone
A bridge is a useful way to see why this matters beyond any single paper. An engineer predicts a bridge will stand for fifty years, on the basis of claims like "this bolt can hold X weight." We don't find out if that claim was true until fifty years have passed, or until it fails. We obviously can't wait that long, or build a thousand test bridges, before deciding whether to build the real one, so instead we substitute validity: how rigorously was the bolt tested, under what conditions, against what standard?
But notice that "this bolt can hold X weight" itself rests on an underlying materials-science claim, which rests on something more basic still. Claims are not independent facts sitting in a flat list. They form a network; each one resting on the claims beneath it, and, once accepted, becoming something that later claims get built on top of.
That structure creates what I'd call dependency risk. If a claim's validity is later undermined, the weakness doesn't stay where it started. It propagates upward into everything built on it. A crack low in that chain threatens the structure above it, regardless of how sound the rest of the structure is.
Citation is the visible trace of this network. It’s one claim pointing at another and saying, in effect, I depend on you. A claim's citation count is a rough proxy for how load-bearing it has become: how much of the field is now relying on it holding.
Why the paper is the wrong unit
This is also, to me, the clearest argument for why evaluating "the paper" was always a slightly blunt instrument. A single paper commonly advances five or ten distinct claims. Nine might rest on careful, well-controlled evidence. One might be the interpretive leap the authors make in the discussion section, thinly supported, and precisely the sentence that a future paper will cite and build on.
Score the paper as a single unit, with a grade, a venue, or an impact factor, and that distinction disappears. The weakness is still there, still capable of propagating into the literature that depends on it, but nothing about "the paper got into a good journal" tells you which of its ten claims was the shaky one. Evaluation at the level of the claim is not a stylistic preference; it's the only resolution at which the actual risk in the network becomes visible at all.
What this means for how we evaluate science
If truthfulness is largely inaccessible at the moment a claim is made, and validity is our best working substitute, then a way of measuring validity rigorously, and doing it consistently, at the level of the individual claim rather than the paper as a whole is not a shortcut around scientific judgment. It's closer to a formalization of the judgment scientists like me already try, imperfectly, to make every time we read something in our field and ask ourselves which parts we actually believe.
I don't think any tool, human or otherwise, can tell us what's true. That was never really available to us, not at the moment discovery happens anyway. What we can ask for, and what I think we should be building toward, is something that can tell us (quickly, consistently, and at the resolution of the claim rather than the paper) how well supported an idea currently is. That's the best signal we have. It's the one science has always run on. We've simply never had a way to apply it evenly, at scale, across the volume of claims the literature now produces.
researchers
researchers

