Ceci est le complément de notre façon de noter les preuves. Cette page-là explique les niveaux que nous attribuons à chaque affirmation ; celle-ci vous met en main les outils pour vérifier n'importe quelle étude, y compris les nôtres.
Nul besoin de statistiques pour bien le faire. La plupart des erreurs qui comptent viennent d'une poignée de manœuvres faciles à repérer une fois qu'on vous les a montrées.
A study describes what happened to one particular group of people, measured one particular way, by researchers with particular questions and particular funders. Reading one well comes down to four questions: who was in it, what was actually measured, how sure the numbers are, and whether the stated conclusion matches the finding underneath. The same four work on almost everything, from a two-line press release to a 30-page paper.
We grade every claim on this site by one of four labels: Strong, Moderate, Emerging, or Preliminary, laid out in how we grade evidence. The tools below are how you check the study behind any label, ours included.
The Evidence Ladder
The design of a study tells you roughly how far to trust it before you read a single number. Five broad designs sit on a ladder, from the weakest at the bottom to the strongest at the top.
At the bottom is the case report, or a case series of a handful of patients: one person who got a treatment and did well or badly. This is an anecdote. It can raise a question worth studying. It cannot settle one, because there is nobody to compare the patient against.
Next come observational studies, where researchers watch people who already do or do not do a thing. A cross-sectional survey is a snapshot of one moment. A case-control study starts from people who have an outcome and looks backward. A cohort study follows a group forward, sometimes for 10 or 20 years. These can show that two measures rise and fall together, measure how strongly, and reveal a dose-response. They cannot show that one caused the other, because the people who do the thing differ from those who do not in a hundred other ways.
The randomized controlled trial is the tool built to fix confounding. Assigning people to treatment or control by chance makes the two groups the same on average, known factors and unknown ones alike, so a difference afterward can reasonably be laid at the treatment's feet. Blinding, where neither the patient nor the assessor knows who got what, keeps expectation from shaping the result.
At the top, a systematic review gathers every study on a question and appraises them together; a meta-analysis pools their numbers into one estimate. Done well, this is the most reliable thing we have. Done on a pile of weak studies, it is a tidy average of weak studies, no stronger than what went in.
The ladder gives you a starting prior and nothing more. One large, clean cohort can outweigh a small, sloppy randomized trial of 30 people. Design sets your expectation; the methods tell you whether a study meets it.
Correlation Is Not Cause
For years, coffee drinking looked tied to lung cancer, until researchers noticed the coffee drinkers of that era also smoked far more. That is the error behind most bad health headlines, and it takes two forms.
The first is confounding, where a third factor drives both things at once. Ice-cream sales and drownings rise together only because both climb in summer. Hormone therapy is the expensive version: for years it looked heart-protective in observational data, because the women who took it were healthier and wealthier to begin with, a pattern called healthy-user bias. The randomized trials that followed found no protection at all.
The second form is reverse causation, where the arrow points the other way. People in the early stages of an illness often move less, so "sitting is linked to disease" can be the disease causing the sitting, with the true arrow reversed.
Randomization breaks both traps, which is why the ladder puts trials above cohorts. But for many important questions, observational data is all there will ever be, because randomizing people would be cruel or impossible. Four signs then make a causal reading more credible: a clear dose-response, the cause arriving before the effect, a plausible biological mechanism, and the same result turning up across different populations. None is decisive alone. Together they are how careful people reason about causes they cannot randomize.
Relative Versus Absolute Risk
One number misleads more readers than any other on a health page: a risk reduction reported without its baseline. Take a drug that lowers your risk of some event from 2 in 1,000 to 1 in 1,000 over a year. The same result can be written two ways:
- As a relative risk reduction: the drug cuts your risk in half, a 50% reduction.
- As an absolute risk reduction: your risk falls by 1 in 1,000, or 0.1 of a percentage point.
A third way says the same thing more plainly: the number needed to treat, here 1,000. A thousand people take the drug for a year for one of them to avoid the event; the other 999 get only the side effects and the bill. Its mirror is the number needed to harm.
Relative figures are incomplete because they hide the starting risk. A 50% cut off a tiny baseline is a tiny change; a 50% cut off a large baseline is a large one.
When you see a percentage reduction, ask two questions every time: reduction from what baseline, and what is the change in plain counts?
Statistical Versus Real-World Significance
The p-value answers one narrow question, and readers routinely mistake it for a bigger one. If the treatment truly did nothing, how often would a result at least this extreme turn up by chance alone, that is all it reports. A p below 0.05 is a convention, a line somebody drew, not a threshold of truth. It does not give the probability that the finding is real or that the hypothesis holds. A study can clear 0.05 and still be a fluke, especially a small one that measured many things and reported the single result that crossed the line.
The confidence interval is more useful, because it shows the whole range of effects the data are compatible with. A narrow interval means a precise estimate; a wide interval means the study can tell you little. If the interval includes "no effect," the result stays uncertain whatever the p-value reads. Prefer the interval to a bare p every time.
Statistical significance and real-world importance part ways more often than headlines admit. A trial of 50,000 people can pin down a blood-pressure drop of half a point, airtight and useless together, because no patient would feel half a point. A genuinely useful effect can slip the other way, missing significance in a study too small to catch it. That miss reflects the study's size; it does not prove the effect is gone. Ask what size of effect would actually matter to a person, then check whether the study found an effect that big before you check its p-value.
Surrogates Versus Outcomes
Drugs that suppressed the irregular heartbeats after a heart attack were expected to save lives. Tested at last against death itself, they raised it. The heartbeat they fixed was a surrogate, a lab marker standing in for the outcome anyone actually cares about.
Surrogates are everywhere; five common ones are LDL cholesterol for heart attacks, blood pressure for strokes, tumor shrinkage for survival, bone density for fractures, and HbA1c for the complications of diabetes. They are attractive because they move in weeks and cost little, where the real outcomes take years.
Moving the marker does not always move the outcome, and can move it the wrong way. Some cancer drugs shrink a tumor on a scan while never letting people live longer or feel better. So when a study reports a surrogate, treat it as a promising lead, then ask the real question: did anyone measure how the patients did, how long they lived, how they felt, whether the fracture or the heart attack ever came?
Who Was Studied, And Who Paid
A result measured in 2,315 Finnish men is solid for Finnish men and an open question for everyone else. Two studies can report identical numbers and still deserve very different trust, depending entirely on the fine print.
Small studies are noisy. A dramatic result in 12 people is a reason to run a bigger study, not a reason to change what you do.
Who was in the study decides who the study is for. A finding measured in young healthy men, or in one country, or across a narrow band of ages, may not carry to you. Much of exercise and physiology research skews heavily male, and the skew turns invisible once a single number gets quoted and re-quoted. This is why our own claims flag the sex they were measured in. Generalizability is a question you have to ask; a study rarely volunteers its own limits.
A study that declared its main outcome and its analysis plan in advance, then reported exactly that, is far harder to fudge than one that measured 10 things and told you about the two that worked. Look for a preregistered protocol.
The published studies are a biased sample of the studies actually run. Trials that find nothing often go unpublished, so the published record tilts positive before you read a word of it. A systematic review that searched trial registries beyond the journals is the partial cure.
Industry funding is not disqualifying, and plenty of excellent research is paid for by the companies that make the product. It is a reason to read the methods harder, to check who chose the outcomes and the comparison group, and to weigh the result beside independent work.
A result in mice, or in cells in a dish, is a reason to run a human trial and nothing more. Most findings that look striking in a mouse never reproduce in a person.
You can usually trace how a modest paper becomes a loud headline. The headline comes from a press release, the press release from the abstract, and the abstract is often the most flattering part of the entire document.
The Research & Studies
Everything here is based on the research we have collected and checked, sorted into groups and ordered with the strongest evidence first. Click any claim to open the studies behind it.
Evidence And Methods
Les résumés d'environ 58 % des essais non significatifs se lisent encore comme positifs
Des chercheurs ont rassemblé 72 essais dont le résultat principal n'avait pas atteint la significativité statistique, c'est-à-dire que le traitement n'avait pas clairement battu la comparaison. Dans la plupart d'entre eux, le résumé, la partie que presque tout le monde lit, était encore rédigé pour donner l'impression que le traitement fonctionnait.
Boutron et ses collègues ont sélectionné des essais indexés en décembre 2006 et identifié 72 essais randomisés en groupes parallèles avec un critère de jugement principal statistiquement non significatif. Les assesseurs ont classifié le spin, défini comme un rapportage qui met l'accent sur le bénéfice du traitement malgré le résultat non significatif. Le spin est apparu dans la section des résultats du résumé dans 37.5 % des rapports et dans la section des conclusions du résumé dans 58.3 %, les conclusions du texte principal étant affectées à des taux similaires ou plus élevés. Le résumé est la partie que la plupart des lecteurs voient, et souvent la seule, donc le formuler de manière optimiste façonne la façon dont un essai nul est mémorisé.
Who this may not transfer to:This is a finding about how trials get written up, not a health effect, so it applies whenever you read a study, not to any one group of people.
Lire le critère de jugement principal et son résultat avant de lire la conclusion. Si le critère principal prédéfini n'a pas bougé mais que la conclusion est positive, cette phrase s'appuie généralement sur un résultat secondaire ou un sous-groupe. Se fier aux chiffres, pas aux adjectifs.
The study · 1
Boutron et al., reporting and interpretation of randomized trials with statistically nonsignificant results for primary outcomes · JAMA 2010;303(20):2058-64
Les revues montraient 94 % des essais sur les antidépresseurs positifs ; les données de la FDA en montraient 51 %
Les autorités réglementaires détenaient les résultats de 74 essais sur les antidépresseurs. Presque chaque essai avec un bon résultat a été publié, tandis que la plupart des essais avec un mauvais résultat ne l'ont jamais été, ou ont été rédigés pour sembler meilleurs qu'ils n'étaient. Quelqu'un qui ne lit que les revues scientifiques penserait que presque tous les essais ont réussi ; le bilan complet montrait qu'environ la moitié l'avaient fait.
Turner et ses collègues ont comparé 74 essais sur les antidépresseurs enregistrés auprès de la Food and Drug Administration américaine entre 1987 et 2004 avec ce qui a atteint la littérature publiée. Sur 38 études que la FDA a jugées positives, 37 ont été publiées. Sur 36 études avec des résultats négatifs ou discutables, 22 n'ont pas été publiées du tout et 11 ont été publiées sous une forme qui transmettait un résultat positif. Le bilan publié montrait donc 94 % des essais comme positifs contre 51 % selon la FDA, et les tailles d'effet dans la littérature des revues étaient environ 32 % plus grandes que celles du dataset complet de la FDA, avec des médicaments individuels gonflés entre 11 % et 69 %.
Who this may not transfer to:This is a finding about which trials reach the published record, not a health effect, so it applies whenever you rely on the literature, not to any one group of people.
Un seul essai positif publié ne règle pas grand-chose en soi, car les essais qui ont échoué peuvent simplement être manquants. Préférer une revue systématique qui a consulté les registres d'essais, pas seulement les revues scientifiques, et vérifier si le critère de jugement a été préenregistré avant le début de l'étude.
The study · 1
Turner et al., selective publication of antidepressant trials and its influence on apparent efficacy · N Engl J Med 2008;358(3):252-60
Le cadrage en risque relatif fait paraître le même bénéfice plus grand qu'en risque absolu
Dire aux gens qu'un médicament réduit leur risque de moitié les rend plus enclins à le prendre que si vous leur dites qu'il réduit leur risque de 2 sur 1,000 à 1 sur 1,000, même si ce sont les mêmes faits. Le chiffre relatif qui semble plus grand pèse davantage sur la décision.
Covey a regroupé des études expérimentales qui présentaient le même bénéfice de traitement aux participants sous différents formats numériques. La réduction du risque relatif produisait constamment des évaluations d'efficacité plus élevées et une plus grande disposition à prendre ou recommander un traitement que la réduction du risque absolu ou le nombre nécessaire à traiter, chez les participants profanes, les étudiants et les cliniciens. Le résultat concerne la façon dont un format oriente la perception, pas lequel est correct : le chiffre absolu et le nombre nécessaire à traiter portent le risque de base que le chiffre relatif omet.
Who this may not transfer to:This is a finding about how risk numbers steer perception, not a health effect, so it applies whenever you meet a percentage reduction, not to any one group of people.
Quand vous rencontrez une réduction en pourcentage, posez deux questions à chaque fois : réduction à partir de quel risque de départ, et quelle est la variation en chiffres simples. Si le risque de base est faible, une grande réduction relative peut n'être qu'une réduction absolue triviale.
The study · 1
Covey, a meta-analysis of the effects of presenting treatment benefits in different formats · Med Decis Making 2007;27(5):638-54
Questions To Ask Any Study
Run any study, headline, or press release through the same seven questions. Ninety seconds is enough to screen it, and the answers tell you whether the finding earns a closer look.
Common Questions
What is the single most useful question to ask about a study?
Whether the group studied resembles you, and whether what they measured is what you care about. A large effect on a surrogate marker in people nothing like you guides your decision less than a modest effect on a real outcome in people who share your situation. Relevance comes before design and statistics.
Why is "cuts your risk by half" misleading?
Because it hides the starting risk. Halving a risk of 2 in 1,000 saves one person for every 1,000 treated; halving a risk of 1 in 3 is enormous. The relative figure reads the same in both cases and tells you almost nothing alone. Convert it to the absolute change, and where you can, the number needed to treat.
Does one study ever settle anything?
Rarely. A finding earns trust when it replicates across different groups and different research teams. A good systematic review checks exactly that. A single trial, however large, is one data point. Be most wary of a lone small study that overturns a large body of prior work; that is more often noise than revolution.
Is a study in mice or cells worthless?
No, but it answers a different question. Animal and laboratory work is how researchers find mechanisms and decide what is worth testing in people. It points toward a human trial worth running. Most striking results in mice do not reproduce in humans, so read "in mice" in a headline as a signal that the human evidence is not in yet.
How can I tell a real finding from an overblown headline?
Look for the tells of a claim reaching past its evidence. "Linked to" or "associated with" usually marks observational data that cannot establish cause. A relative figure arrives with no absolute one beside it. A surrogate marker gets reported as though it were an outcome. The sample is small. And "may" carries the whole first sentence. When several tells stack up in one headline, go find the study and read the primary outcome yourself.
Explore Related
Other pages this one connects to, by the evidence they share, the outcomes they touch, and the ground they cover.
All 3 sources on this page independently checked and cross-referenced.
Thomas Dehli, Founder & Editor, Sacred Lotus
Sacred Lotus has published Chinese medicine reference material since 2001. Integrative pages are held to the same standard as the herb and formula library: cite the source, grade the claim at its real strength, and say where the research has not looked. This page is educational and it is not medical advice. Last reviewed and updated August 10, 2026.