This is the companion to how we grade evidence. That page explains the tiers we put on every claim; this one hands you the tools to check any study, including ours.
You do not need statistics to do it well. Most of the mistakes that matter come from a handful of moves that are easy to spot once someone points them out.
A study describes what happened to a particular group of people, measured a particular way, by researchers with particular questions and particular funders. Reading one well means asking who was in it, what was actually measured, how sure the numbers are, and whether the stated conclusion matches the finding underneath. The same few questions work on almost everything, from a press release to a full paper.
For how we turn all of this into the Strong, Moderate, Emerging and Preliminary labels on our own pages, see how we grade evidence. This page is the reader-facing half of that method.
The Evidence Ladder
Not all studies carry the same weight, and the design tells you roughly how much to trust a finding before you read a single number.
At the bottom sits the case report or case series: one person, or a handful, who got a treatment and did well or badly. This is an anecdote. It can raise a question worth studying, and it cannot settle one, because there is nobody to compare them to.
Above that are observational studies, where researchers watch people who already do or do not do a thing. A cross-sectional survey takes a snapshot. A case-control study starts from people who have an outcome and looks back. A cohort study follows a group forward for years. These can show that two things rise and fall together, measure how strongly, and reveal a dose-response. On their own they cannot show that one caused the other, because the people who do the thing differ from the people who do not in a hundred other ways.
The randomized controlled trial is the tool built to fix that. By assigning people to treatment or control by chance, it makes the two groups the same on average, known factors and unknown ones alike, so a difference afterward can reasonably be attributed to the treatment. Blinding, where neither patient nor assessor knows who got what, keeps the expectations of patient and assessor from shaping the result.
At the top sits the systematic review, which gathers every study on a question and appraises them together, and the meta-analysis, which pools their numbers into one estimate. Done well, this is the most reliable thing we have. Done on a pile of weak studies, it is a tidy average of weak studies, no stronger than what went in.
Treat the ladder as a starting prior, not a verdict. A large, clean cohort can be worth more than a tiny, sloppy randomized trial. The design sets your starting expectation; the methods tell you whether the study meets it.
Correlation Is Not Cause
This is the error that produces most bad health headlines, and it takes two forms.
The first is confounding: a third factor drives both things you are looking at. Ice cream sales and drownings rise together, because both rise in summer. Coffee drinking once looked tied to lung cancer, until you noticed that coffee drinkers of that era smoked more. For years hormone therapy looked heart-protective in observational data, because the women who took it were healthier and better-off to begin with, a pattern called healthy-user bias; the randomized trials that followed did not find the protection.
The second is reverse causation: the arrow points the other way. People with early illness often move less, so "sitting is linked to disease" can partly be disease causing sitting rather than sitting causing disease.
Randomization is what breaks both, which is why the ladder puts trials above cohorts. When only observational data exists, and for many important questions it is all there will ever be, a few things make a causal reading more credible: a clear dose-response, the cause coming before the effect in time, a plausible biological mechanism, and the same result turning up across different populations. None of these is decisive alone. Together they are how careful people reason about causes they cannot randomize.
Relative Versus Absolute Risk
One number misleads more readers than any other: a risk reduction reported without its baseline.
Suppose a drug takes your risk of some event from 2 in 1,000 down to 1 in 1,000 over a year. Two true ways to report that:
- As a relative risk reduction: the drug cuts your risk in half, a 50% reduction.
- As an absolute risk reduction: your risk falls by 1 in 1,000, or 0.1 of a percentage point.
Same trial, same result. The relative figure sounds large; the absolute figure shows how small the change is. A third way states it in plain terms: the number needed to treat, here 1,000, meaning a thousand people take the drug for a year for one of them to avoid the event, and the other 999 get only the side effects and the bill. Its mirror is the number needed to harm.
Relative figures are incomplete because they hide the starting risk. A 50% cut off a tiny baseline is a tiny change; a 50% cut off a large baseline is a large one.
When you see a percentage reduction, ask two questions every time: reduction from what baseline, and what is the change in plain counts?
Statistical Versus Real-World Significance
A p-value answers one narrow question: if the treatment truly did nothing, how often would you see a result at least this extreme just by chance. A p below 0.05 is a convention, a line someone drew, not a truth threshold, and it is not the probability that the finding is real or that the hypothesis is true. A study can clear it and still be a fluke, especially a small study that measured many things and reported the one that crossed the line.
A confidence interval is more useful, because it shows the whole range of effects the data are compatible with. A narrow interval means a precise estimate; a wide one means the study cannot tell you much. If the interval includes "no effect", the result is uncertain whatever the p-value says. Prefer the interval to a bare p every time.
Statistical significance and real-world importance are different things. A trial with fifty thousand people can find a difference that is statistically rock-solid and clinically pointless, a blood-pressure drop of half a point that no patient would ever feel. And a useful effect can miss significance in a study too small to detect it, which is a failure of the study, not evidence the effect is absent. Ask what size of effect would actually matter to a person, then see whether the study found one, not just whether it found a p-value.
Surrogates Versus Outcomes
A surrogate endpoint is a lab marker that stands in for the thing you actually care about: LDL cholesterol for heart attacks, blood pressure for strokes, tumor shrinkage for survival, bone density for fractures, HbA1c for the complications of diabetes. Surrogates are attractive because they move fast and cheap, where the real outcomes take years.
The trouble is that moving the marker does not always move the outcome, and can move it the wrong way. Antiarrhythmic drugs that suppressed the irregular heartbeats after a heart attack were found, when finally tested against death, to increase it. Some cancer drugs delay tumor growth on a scan while never letting people live longer or better. When a study reports a surrogate, treat it as a promising lead and ask the next question: did anyone measure how the patients actually did, how long they lived, how they felt, whether the fracture or the heart attack came?
Who Was Studied, And Who Paid
Two studies with identical numbers can deserve very different trust, depending on the fine print.
Sample size. Small studies are noisy. A dramatic result in twelve people is a reason to run a bigger study, not a reason to change what you do.
Who was in it. A finding measured in young healthy men, or in one country, or in a narrow age band, may not carry to you. Much of exercise and physiology research skews heavily male, and the skew becomes invisible once a number gets repeated. A result in 2,315 Finnish men is real for Finnish men and an open question for everyone else. This is why our own claims flag the sex they were measured in. Generalizability is a question you have to ask; the study rarely volunteers its limits.
Preregistration. A study that declared its main outcome and analysis plan in advance, and then reported that, is far harder to fudge than one that measured ten things and told you about the two that worked. Look for a preregistered protocol.
Publication bias. The studies that get published are not a fair sample of the studies that get run. Trials that find nothing often go unpublished, so the published record skews positive before anyone reads a word of it. A systematic review that searched trial registries, not just journals, is the partial cure.
Funding and conflicts. Industry funding is not disqualifying, and plenty of excellent research is paid for by the companies that make the product. It is a reason to read the methods harder, to check who chose the outcomes and the comparison, and to weight the result alongside independent work.
Animals and mechanisms. A result in mice, or in cells in a dish, is a reason to run a human trial, not a reason to act. Most findings that look striking in a mouse never reproduce in a person.
Put those together and you can usually see how a modest paper becomes a loud headline. The headline comes from a press release, the press release from the abstract, and the abstract, as the research below shows, is often the most flattering part of the whole document.
The Research & Studies
Everything here is based on the research we have collected and checked, sorted into groups and ordered with the strongest evidence first. Click any claim to open the studies behind it.
Evidence And Methods
Abstracts of about 58% of nonsignificant trials still read as positive
Researchers gathered 72 trials where the main result did not reach statistical significance, meaning the treatment had not clearly beaten the comparison. In most of them the abstract, the part almost everyone reads, was still written to sound like the treatment worked.
Boutron and colleagues screened trials indexed in December 2006 and identified 72 parallel-group randomized trials with a statistically nonsignificant primary outcome. Assessors classified spin, defined as reporting that emphasises the treatment's benefit despite the nonsignificant result. Spin appeared in the results section of the abstract in 37.5% of the reports and in the conclusions section of the abstract in 58.3%, with the main-text conclusions affected at similar or higher rates. The abstract is the part most readers see, and often the only part, so wording it optimistically shapes how a null trial is remembered.
Who this may not transfer to:This is a finding about how trials get written up, not a health effect, so it applies whenever you read a study rather than to any one group of people.
Read the primary outcome and its result before you read the conclusion. If the pre-specified main endpoint did not move but the conclusion is upbeat, that sentence is usually leaning on a secondary finding or a subgroup. Trust the number, not the adjectives.
The study · 1
Boutron et al., reporting and interpretation of randomized trials with statistically nonsignificant results for primary outcomes · JAMA 2010;303(20):2058-64
Journals showed 94% of antidepressant trials positive; the FDA's data showed 51%
Regulators held the results of 74 antidepressant trials. Almost every trial with a good result got published, while most trials with a poor result never did, or were written up to look better than they were. Someone reading only the journals would think nearly every trial succeeded; the full record showed about half did.
Turner and colleagues compared 74 antidepressant trials registered with the US Food and Drug Administration between 1987 and 2004 against what reached the published literature. Of 38 studies the FDA judged positive, 37 were published. Of 36 studies with negative or questionable results, 22 were not published at all and 11 were published in a form that conveyed a positive outcome. The published record therefore showed 94% of trials as positive against the FDA's 51%, and the effect sizes in the journal literature ran about 32% larger than those from the full FDA dataset, with individual drugs inflated between 11% and 69%.
Who this may not transfer to:This is a finding about which trials reach the published record, not a health effect, so it applies whenever you rely on the literature rather than to any one group of people.
One published positive trial settles little on its own, because the trials that failed may simply be missing. Prefer a systematic review that searched trial registries, not journals alone, and check whether the outcome was preregistered before the study began.
The study · 1
Turner et al., selective publication of antidepressant trials and its influence on apparent efficacy · N Engl J Med 2008;358(3):252-60
Relative-risk framing makes the same benefit look bigger than absolute risk
Tell people a pill cuts their risk in half and they are keener on it than if you tell them it takes their risk from 2 in 1,000 to 1 in 1,000, even though those are the same fact. The bigger-sounding relative number pulls harder on the decision.
Covey pooled experimental studies that presented the same treatment benefit to participants in different numerical formats. Relative risk reduction consistently produced higher ratings of effectiveness and greater willingness to take or recommend a treatment than absolute risk reduction or the number needed to treat, across lay participants, students and clinicians. The finding is about how a format steers perception, not about which format is correct: the absolute figure and the number needed to treat carry the baseline risk that the relative figure leaves out.
Who this may not transfer to:This is a finding about how risk numbers steer perception, not a health effect, so it applies whenever you meet a percentage reduction rather than to any one group of people.
When you meet a percentage reduction, ask two things every time: reduction from what starting risk, and what is the change in plain counts. If the baseline risk is small, a large relative reduction can be a trivial absolute one.
The study · 1
Covey, a meta-analysis of the effects of presenting treatment benefits in different formats · Med Decis Making 2007;27(5):638-54
Questions To Ask Any Study
You do not need to do this for every article you skim. Run it on the ones that might change what you do.
Common Questions
What is the single most useful question to ask about a study?
Whether the group studied resembles you, and whether the thing they measured is the thing you care about. A large effect on a surrogate marker in people nothing like you is a weaker guide to your own decision than a modest effect on a real outcome in people who share your situation. Design and statistics matter, but relevance comes first.
Why is "cuts your risk by half" misleading?
Because it hides the starting risk. Halving a risk of 2 in 1,000 saves one person per thousand treated; halving a risk of 1 in 3 is enormous. The relative figure is the same in both cases and tells you almost nothing on its own. Always convert it to the absolute change and, where you can, the number needed to treat.
Does one study ever settle anything?
Rarely. Findings become trustworthy when they replicate across different groups and different research teams, which is exactly what a good systematic review checks. A single trial, however large, is one data point. Be especially wary of a lone small study that overturns a large body of prior work; it is more often noise than revolution.
Is a study in mice or cells worthless?
No, but it answers a different question. Animal and laboratory work is how researchers find mechanisms and decide what is worth testing in people. It is a reason to run a human trial, not a reason to act. Most striking results in mice do not reproduce in humans, so treat "in mice" in a headline as a flag that the human evidence is not in yet.
How can I tell a real finding from an overblown headline?
Look for the tells of a claim leaning past its evidence: "linked to" or "associated with", which usually signal observational data that cannot establish cause; a relative figure with no absolute one; a surrogate marker reported as if it were an outcome; a small sample; and "may" doing heavy lifting in the first sentence. When several of those stack up in one headline, go find the study and read the primary outcome yourself.
Explore Related
Other pages this one connects to, by the evidence they share, the outcomes they touch, and the ground they cover.
All 3 sources on this page independently checked and cross-referenced.
Thomas Dehli, Founder & Editor, Sacred Lotus
Sacred Lotus has published Chinese medicine reference material since 2001. Integrative pages are held to the same standard as the herb and formula library: cite the source, grade the claim at its real strength, and say where the research has not looked. This page is educational and it is not medical advice. Last reviewed and updated August 10, 2026.