Dies ist die Ergänzung zu unserer Methode der Evidenzbewertung. Die andere Seite erklärt die Stufen, die wir jeder Aussage zuordnen; diese hier gibt Ihnen die Werkzeuge an die Hand, um jede Studie – einschließlich unserer eigenen – zu überprüfen.
Sie benötigen keine Statistikkenntnisse, um dies gut zu machen. Die meisten relevanten Fehler stammen von einer kleinen Anzahl von Vorgehensweisen, die leicht erkennbar sind, sobald jemand darauf hinweist.
A study describes what happened to one particular group of people, measured one particular way, by researchers with particular questions and particular funders. Reading one well comes down to four questions: who was in it, what was actually measured, how sure the numbers are, and whether the stated conclusion matches the finding underneath. The same four work on almost everything, from a two-line press release to a 30-page paper.
We grade every claim on this site by one of four labels: Strong, Moderate, Emerging, or Preliminary, laid out in how we grade evidence. The tools below are how you check the study behind any label, ours included.
The Evidence Ladder
The design of a study tells you roughly how far to trust it before you read a single number. Five broad designs sit on a ladder, from the weakest at the bottom to the strongest at the top.
At the bottom is the case report, or a case series of a handful of patients: one person who got a treatment and did well or badly. This is an anecdote. It can raise a question worth studying. It cannot settle one, because there is nobody to compare the patient against.
Next come observational studies, where researchers watch people who already do or do not do a thing. A cross-sectional survey is a snapshot of one moment. A case-control study starts from people who have an outcome and looks backward. A cohort study follows a group forward, sometimes for 10 or 20 years. These can show that two measures rise and fall together, measure how strongly, and reveal a dose-response. They cannot show that one caused the other, because the people who do the thing differ from those who do not in a hundred other ways.
The randomized controlled trial is the tool built to fix confounding. Assigning people to treatment or control by chance makes the two groups the same on average, known factors and unknown ones alike, so a difference afterward can reasonably be laid at the treatment's feet. Blinding, where neither the patient nor the assessor knows who got what, keeps expectation from shaping the result.
At the top, a systematic review gathers every study on a question and appraises them together; a meta-analysis pools their numbers into one estimate. Done well, this is the most reliable thing we have. Done on a pile of weak studies, it is a tidy average of weak studies, no stronger than what went in.
The ladder gives you a starting prior and nothing more. One large, clean cohort can outweigh a small, sloppy randomized trial of 30 people. Design sets your expectation; the methods tell you whether a study meets it.
Correlation Is Not Cause
For years, coffee drinking looked tied to lung cancer, until researchers noticed the coffee drinkers of that era also smoked far more. That is the error behind most bad health headlines, and it takes two forms.
The first is confounding, where a third factor drives both things at once. Ice-cream sales and drownings rise together only because both climb in summer. Hormone therapy is the expensive version: for years it looked heart-protective in observational data, because the women who took it were healthier and wealthier to begin with, a pattern called healthy-user bias. The randomized trials that followed found no protection at all.
The second form is reverse causation, where the arrow points the other way. People in the early stages of an illness often move less, so "sitting is linked to disease" can be the disease causing the sitting, with the true arrow reversed.
Randomization breaks both traps, which is why the ladder puts trials above cohorts. But for many important questions, observational data is all there will ever be, because randomizing people would be cruel or impossible. Four signs then make a causal reading more credible: a clear dose-response, the cause arriving before the effect, a plausible biological mechanism, and the same result turning up across different populations. None is decisive alone. Together they are how careful people reason about causes they cannot randomize.
Relative Versus Absolute Risk
One number misleads more readers than any other on a health page: a risk reduction reported without its baseline. Take a drug that lowers your risk of some event from 2 in 1,000 to 1 in 1,000 over a year. The same result can be written two ways:
- As a relative risk reduction: the drug cuts your risk in half, a 50% reduction.
- As an absolute risk reduction: your risk falls by 1 in 1,000, or 0.1 of a percentage point.
A third way says the same thing more plainly: the number needed to treat, here 1,000. A thousand people take the drug for a year for one of them to avoid the event; the other 999 get only the side effects and the bill. Its mirror is the number needed to harm.
Relative figures are incomplete because they hide the starting risk. A 50% cut off a tiny baseline is a tiny change; a 50% cut off a large baseline is a large one.
When you see a percentage reduction, ask two questions every time: reduction from what baseline, and what is the change in plain counts?
Statistical Versus Real-World Significance
The p-value answers one narrow question, and readers routinely mistake it for a bigger one. If the treatment truly did nothing, how often would a result at least this extreme turn up by chance alone, that is all it reports. A p below 0.05 is a convention, a line somebody drew, not a threshold of truth. It does not give the probability that the finding is real or that the hypothesis holds. A study can clear 0.05 and still be a fluke, especially a small one that measured many things and reported the single result that crossed the line.
The confidence interval is more useful, because it shows the whole range of effects the data are compatible with. A narrow interval means a precise estimate; a wide interval means the study can tell you little. If the interval includes "no effect," the result stays uncertain whatever the p-value reads. Prefer the interval to a bare p every time.
Statistical significance and real-world importance part ways more often than headlines admit. A trial of 50,000 people can pin down a blood-pressure drop of half a point, airtight and useless together, because no patient would feel half a point. A genuinely useful effect can slip the other way, missing significance in a study too small to catch it. That miss reflects the study's size; it does not prove the effect is gone. Ask what size of effect would actually matter to a person, then check whether the study found an effect that big before you check its p-value.
Surrogates Versus Outcomes
Drugs that suppressed the irregular heartbeats after a heart attack were expected to save lives. Tested at last against death itself, they raised it. The heartbeat they fixed was a surrogate, a lab marker standing in for the outcome anyone actually cares about.
Surrogates are everywhere; five common ones are LDL cholesterol for heart attacks, blood pressure for strokes, tumor shrinkage for survival, bone density for fractures, and HbA1c for the complications of diabetes. They are attractive because they move in weeks and cost little, where the real outcomes take years.
Moving the marker does not always move the outcome, and can move it the wrong way. Some cancer drugs shrink a tumor on a scan while never letting people live longer or feel better. So when a study reports a surrogate, treat it as a promising lead, then ask the real question: did anyone measure how the patients did, how long they lived, how they felt, whether the fracture or the heart attack ever came?
Who Was Studied, And Who Paid
A result measured in 2,315 Finnish men is solid for Finnish men and an open question for everyone else. Two studies can report identical numbers and still deserve very different trust, depending entirely on the fine print.
Small studies are noisy. A dramatic result in 12 people is a reason to run a bigger study, not a reason to change what you do.
Who was in the study decides who the study is for. A finding measured in young healthy men, or in one country, or across a narrow band of ages, may not carry to you. Much of exercise and physiology research skews heavily male, and the skew turns invisible once a single number gets quoted and re-quoted. This is why our own claims flag the sex they were measured in. Generalizability is a question you have to ask; a study rarely volunteers its own limits.
A study that declared its main outcome and its analysis plan in advance, then reported exactly that, is far harder to fudge than one that measured 10 things and told you about the two that worked. Look for a preregistered protocol.
The published studies are a biased sample of the studies actually run. Trials that find nothing often go unpublished, so the published record tilts positive before you read a word of it. A systematic review that searched trial registries beyond the journals is the partial cure.
Industry funding is not disqualifying, and plenty of excellent research is paid for by the companies that make the product. It is a reason to read the methods harder, to check who chose the outcomes and the comparison group, and to weigh the result beside independent work.
A result in mice, or in cells in a dish, is a reason to run a human trial and nothing more. Most findings that look striking in a mouse never reproduce in a person.
You can usually trace how a modest paper becomes a loud headline. The headline comes from a press release, the press release from the abstract, and the abstract is often the most flattering part of the entire document.
The Research & Studies
Everything here is based on the research we have collected and checked, sorted into groups and ordered with the strongest evidence first. Click any claim to open the studies behind it.
Evidence And Methods
Abstracts of about 58% of nonsignificant trials still read as positive
Forschende sammelten 72 Studien, bei denen das Hauptergebnis keine statistische Signifikanz erreichte, das heißt, die Behandlung hatte die Vergleichsgruppe nicht eindeutig übertroffen. Bei den meisten davon war der Abstract, der Teil, den fast jeder liest, dennoch so formuliert, dass es klang, als hätte die Behandlung gewirkt.
Boutron and colleagues screened trials indexed in December 2006 and identified 72 parallel-group randomized trials with a statistically nonsignificant primary outcome. Assessors classified spin, defined as reporting that emphasises the treatment's benefit despite the nonsignificant result. Spin appeared in the results section of the abstract in 37.5% of the reports and in the conclusions section of the abstract in 58.3%, with the main-text conclusions affected at similar or higher rates. The abstract is the part most readers see, and often the only part, so wording it optimistically shapes how a null trial is remembered.
Who this may not transfer to:This is a finding about how trials get written up, not a health effect, so it applies whenever you read a study, not to any one group of people.
Read the primary outcome and its result before you read the conclusion. If the pre-specified main endpoint did not move but the conclusion is upbeat, that sentence is usually leaning on a secondary finding or a subgroup. Trust the number, not the adjectives.
The study · 1
Boutron et al., reporting and interpretation of randomized trials with statistically nonsignificant results for primary outcomes · JAMA 2010;303(20):2058-64
Journals showed 94% of antidepressant trials positive; the FDA's data showed 51%
Regulators held the results of 74 antidepressant trials. Almost every trial with a good result got published, while most trials with a poor result never did, or were written up to look better than they were. Someone reading only the journals would think nearly every trial succeeded; the full record showed about half did.
Turner and colleagues compared 74 antidepressant trials registered with the US Food and Drug Administration between 1987 and 2004 against what reached the published literature. Of 38 studies the FDA judged positive, 37 were published. Of 36 studies with negative or questionable results, 22 were not published at all and 11 were published in a form that conveyed a positive outcome. The published record therefore showed 94% of trials as positive against the FDA's 51%, and the effect sizes in the journal literature ran about 32% larger than those from the full FDA dataset, with individual drugs inflated between 11% and 69%.
Who this may not transfer to:This is a finding about which trials reach the published record, not a health effect, so it applies whenever you rely on the literature, not to any one group of people.
One published positive trial settles little on its own, because the trials that failed may simply be missing. Prefer a systematic review that searched trial registries, not journals alone, and check whether the outcome was preregistered before the study began.
The study · 1
Turner et al., selective publication of antidepressant trials and its influence on apparent efficacy · N Engl J Med 2008;358(3):252-60
Relative-risk framing makes the same benefit look bigger than absolute risk
Tell people a pill cuts their risk in half and they are keener on it than if you tell them it takes their risk from 2 in 1,000 to 1 in 1,000, even though those are the same fact. The bigger-sounding relative number pulls harder on the decision.
Covey pooled experimental studies that presented the same treatment benefit to participants in different numerical formats. Relative risk reduction consistently produced higher ratings of effectiveness and greater willingness to take or recommend a treatment than absolute risk reduction or the number needed to treat, across lay participants, students and clinicians. The finding is about how a format steers perception, not about which format is correct: the absolute figure and the number needed to treat carry the baseline risk that the relative figure leaves out.
Who this may not transfer to:This is a finding about how risk numbers steer perception, not a health effect, so it applies whenever you meet a percentage reduction, not to any one group of people.
When you meet a percentage reduction, ask two things every time: reduction from what starting risk, and what is the change in plain counts. If the baseline risk is small, a large relative reduction can be a trivial absolute one.
The study · 1
Covey, a meta-analysis of the effects of presenting treatment benefits in different formats · Med Decis Making 2007;27(5):638-54
Questions To Ask Any Study
Run any study, headline, or press release through the same seven questions. Ninety seconds is enough to screen it, and the answers tell you whether the finding earns a closer look.
Common Questions
What is the single most useful question to ask about a study?
Whether the group studied resembles you, and whether what they measured is what you care about. A large effect on a surrogate marker in people nothing like you guides your decision less than a modest effect on a real outcome in people who share your situation. Relevance comes before design and statistics.
Why is "cuts your risk by half" misleading?
Because it hides the starting risk. Halving a risk of 2 in 1,000 saves one person for every 1,000 treated; halving a risk of 1 in 3 is enormous. The relative figure reads the same in both cases and tells you almost nothing alone. Convert it to the absolute change, and where you can, the number needed to treat.
Does one study ever settle anything?
Rarely. A finding earns trust when it replicates across different groups and different research teams. A good systematic review checks exactly that. A single trial, however large, is one data point. Be most wary of a lone small study that overturns a large body of prior work; that is more often noise than revolution.
Is a study in mice or cells worthless?
No, but it answers a different question. Animal and laboratory work is how researchers find mechanisms and decide what is worth testing in people. It points toward a human trial worth running. Most striking results in mice do not reproduce in humans, so read "in mice" in a headline as a signal that the human evidence is not in yet.
How can I tell a real finding from an overblown headline?
Look for the tells of a claim reaching past its evidence. "Linked to" or "associated with" usually marks observational data that cannot establish cause. A relative figure arrives with no absolute one beside it. A surrogate marker gets reported as though it were an outcome. The sample is small. And "may" carries the whole first sentence. When several tells stack up in one headline, go find the study and read the primary outcome yourself.
Explore Related
Other pages this one connects to, by the evidence they share, the outcomes they touch, and the ground they cover.
All 3 sources on this page independently checked and cross-referenced.
Thomas Dehli, Founder & Editor, Sacred Lotus
Sacred Lotus has published Chinese medicine reference material since 2001. Integrative pages are held to the same standard as the herb and formula library: cite the source, grade the claim at its real strength, and say where the research has not looked. This page is educational and it is not medical advice. Last reviewed and updated August 10, 2026.