A chair, a wall, a stopwatch and a well-fitted blood pressure cuff will tell you more about the next ten years than most panels of blood tests, and most of it costs nothing.
This is a guide to the body numbers you can track yourself:
- the ones that predict health
- how to measure each without fooling yourself
- which ones consumer gadgets get right and which they get wrong
- where the limits lie
Grip strength and walking speed carry the strongest links to survival, and both are readouts rather than levers, so training the test moves the score without moving what the score predicted. Blood pressure is the exception worth the most attention, because measuring it badly adds more error than most changes you could make in response, and a week of home readings predicts events better than a single reading in a clinic.
The limits matter too: a proper fitness test needs a lab, screening yourself for everything does not lower mortality, and continuous tracking can feed anxiety.
Findings & Outcomes
What It Is
Most of what predicts how the next decade goes can be measured in your own kitchen in about fifteen minutes, without a blood draw, and the tests with the largest associations attached to them are the ones that cost nothing. This page is the overview of those measures. Each one links out to its own deeper page where there is one, and each arrives here with the protocol its published numbers came from, what the score predicts, and what it does not.
One distinction runs through the whole page, and it decides what to do with a bad score: the difference between markers and levers. A marker reads your current state. A lever changes it. Practicing a balance test makes you better at the balance test, and whether that changes the risk the test was reading is a separate question with a separate literature, thin or absent for most of these measures. Grip strength and the balance tests are markers you cannot usefully train toward. Walking speed, the chair stand and blood pressure are the ones that respond to being acted on. The sections below say which is which as they go.
Strength, Balance and Walking
These are the cheapest measures and they carry the strongest links to survival, because a squeeze, a stand and a walk each draw on muscle, nerve, heart and lung at once, so the number reads the whole body rather than the part being tested.
Grip strength is the clearest marker. Across 139,691 adults in 17 countries, each 11 lb (5 kg) less grip came with a 16% higher rate of death. Squeezing a tennis ball daily raises your grip number with no sign it touches the mortality curve, because the reading stands in for the condition of the whole body rather than the hands. In the same cohort grip showed no association at all with injury from a fall, with fracture, or with new diabetes: a good general alarm and a poor specific one. Getting a number comparable to the published one needs a hand dynamometer, and the full protocol, the age and sex reference ranges, and how to raise it are on the grip strength page.
Walking speed is the most under-used measure here, and it behaves differently from a fixed marker. Speed predicts survival across its whole range, and improvement in it predicts survival too. The pooled survival analysis was built on a stopwatch and a 13-foot (4-meter) course.
-
Mark a 13-foot (4-meter) course
Tape or two objects, 13 feet (4 meters) apart, on a flat floor with a couple of meters of run-off at each end. This is the course the pooled survival data was built on.
One measured 13-foot (4 m) stretch. The pooled cohort of 34,485 adults over 65 averaged 0.92 m/s.
-
Walk it at your usual pace
From a standing start, walk at your usual pace, as if walking down the street. No encouragement, no target, no counting out loud. Normal shoes, and your walking aid if you use one.
Time from first movement to crossing the far mark. Take the better of two walks.
-
Divide 4 by your time in seconds
Four meters in 5 seconds is 0.8 m/s. Four meters in 4 seconds is 1.0 m/s.
Every 0.1 m/s faster came with about 12% lower mortality across the 34,485 participants.
In 439 adults over 65 followed eight years, of six health and function measures only improved gait speed predicted survival: mortality was 31.6% in those who improved by at least 0.1 m/s against 49.3% in those who never did. That earns walking speed the training attention grip does not, and the intervention with a measured effect on it is resistance training, which raised gait speed by 0.08 m/s across 24 trials. What each speed band means for your age and sex, and how to move it, are on the walking speed page, and walking covers what the step-count evidence does and does not support.
The 30-second chair stand reads lower-body strength, the capacity that lets you recover from a stumble and get off the floor afterwards.
-
Set the chair
A straight-backed chair without arms, seat height 17 inches (43.2 cm), placed against a wall so it cannot slide. Sit in the middle, back straight, feet about shoulder width apart and flat.
Seat height is part of the test. A lower sofa or a higher stool gives a number that compares to nothing published.
-
Cross your arms
Arms crossed at the wrists against the chest, and they stay there for the whole test. Pushing off your thighs measures something else.
Needing your arms to stand at all scores zero, and zero is a result. Record it and stop.
-
Stand fully and sit fully, for 30 seconds
A full stand and full sit counts as one, sitting all the way down between each. Count complete stands in 30 seconds.
The count tracked weight-adjusted leg press at r = 0.71 to 0.78 in the original validation.
The US Centers for Disease Control uses the chair stand inside its STEADI fall-risk screen, flagging a count below the age and sex average. Those cut-offs describe the capacity associated with staying physically independent rather than a study of falls, and the full table is on the balance and falls page. This one responds to training more than any other test here: progressive resistance training produced a moderate to large improvement in getting out of a chair. Whether the improved score carries the fall risk with it is the part nobody has shown.
The 10-second one-legged stand is the most shared statistic in this field, and the one most often stripped of its conditions. Stand barefoot beside a counter you can grab, place the free foot behind the standing calf, and time until you touch down or grab hold, with up to three tries on either leg. Ten seconds is the threshold the outcome data was collected against. In 1,702 adults aged 51 to 75 at a Brazilian exercise-medicine clinic, one in five could not hold it, and 17.5% of those people died over a median seven years against 4.6% of those who could, an adjusted hazard ratio of 1.84. It is a self-selected clinic sample, 68% men, with a wide confidence interval, so read it as a free signal worth acting on rather than a precise multiplier, and compare yourself to your own age band, since failing at 54 means something different from failing at 74.
The sitting-rising test measures strength and flexibility together, from how many hand or knee supports you need to lower yourself to the floor and stand back up, scored 0 to 10. In 2,002 adults from the same Brazilian clinic, the lowest scorers (0 to 3) carried 5.44 times the death rate of a top score, and each one-point rise carried a 21% improvement in survival. It shares that clinic population's limits, and because it asks you onto the floor it is not one to attempt alone if you are frail or have fallen in the past year.
The timed up and go is the test a clinic is most likely to hand you, and its numbers need reading carefully. It measures mobility reliably, correlating with the Berg Balance Scale and with an activities-of-daily-living index. As a fall predictor it is weak: pooled across 10 studies, the 13.5-second threshold had specificity 0.74 and sensitivity 0.31, so it misses roughly two thirds of the people who go on to fall, and the review's own regression could not show the score predicts falls at all. That is the reason this page gives no table of tidy thresholds. A cut-off with a sensitivity of 0.31 sorts almost nobody correctly, and printing it beside a number like the grip hazard ratio would imply the two are the same kind of statement.
Blood Pressure, Done Right
This is the most useful thing on the page.
The error introduced by measuring badly is larger than the effect of almost any change you could make in response to the reading.
Against a correctly sized cuff, a regular cuff on an arm that needs an extra-large one read 19.5 mmHg higher systolic in a randomized crossover trial, and 4.8 mmHg higher on an arm needing a large cuff. Against an arm supported on a desk at heart height, an unsupported arm hanging at the side read 6.5 mmHg higher and a hand resting in the lap 3.9 mmHg higher. Stack a wrong cuff on an unsupported arm and technique alone can move a person from normal to a diagnosis.
-
Measure your arm and buy the cuff to fit it
Wrap a tape around the midpoint of your upper arm. The cuff box gives an arm-circumference range. Most cuffs sold with a monitor are the regular size, and a large share of adults need a large or extra-large.
The single largest correctable error: 19.5 mmHg systolic in the people who need an extra-large cuff.
-
Sit still for five minutes first
Back supported, feet flat, legs uncrossed, bladder empty, no coffee or exercise in the previous half hour, and no talking during or between readings.
Five minutes seated before the first reading.
-
Support the arm at heart height
Forearm resting on a table so the middle of the cuff sits level with the middle of your chest. Not held up, not in your lap, not hanging down.
With the arm supported at heart height the reading runs 6.5 mmHg below an unsupported one.
-
Take a week of readings, not one
Two readings each morning and two each evening for seven days, on the same arm, and average the lot. A single reading tells you about that minute.
This is the schedule the trial that showed self-monitoring works actually used.
Home readings earn the effort for a reason beyond convenience. In 2,081 Finnish adults aged 45 to 74 followed nearly seven years, when home and office pressure were entered into the same model, only the home reading still predicted cardiovascular events. The office reading stopped predicting anything.
What the research did, and what we would suggest
The trial dose
TASMINH4 taught 1,182 hypertensive patients to use a validated automated monitor on the non-dominant arm, twice each morning and twice each evening for the first week of every month, and asked their GP to titrate medication on those numbers. At twelve months systolic pressure was 3.5 mmHg lower than usual care with self-monitoring alone, and 4.7 mmHg lower with telemonitoring added.
What we would suggest
If you have been told your pressure is high, do the same week-long block before your appointment and take the whole week's average with you. If you have never been told anything, one properly done week a year is enough to know whether the subject applies to you.
Why the difference: The trial's effect came from the readings reaching a prescriber, not from the readings existing. Self-monitoring on its own, with nobody adjusting anything, has no reason to lower a number. What the week does do is replace a single clinic measurement, taken in the worst conditions available, with an average taken in the conditions the risk associations were established under. Telemonitoring added 1.2 mmHg over self-monitoring alone, which did not reach significance, so the app is not the ingredient.
The smallest version still worth doing: One correctly sized cuff, one arm supported at heart height, one week of mornings. That is most of the available accuracy.
Adjusting it to you: Chinese medicine has no equivalent of a pressure number and does not treat one. What it holds is that the pattern behind a raised reading differs between people: the red-faced, irritable, headachy picture of rising Liver Yang is settled and directed downward, while a pale, exhausted, dizzy-on-standing picture is supported instead. The same figure on the monitor arrives from different places, so the classical reading of it is not the number but everything around it.
The left column is what was measured. The right column is our judgement about what is probably worth doing, which is a different kind of statement, and yours to disagree with.
Blood pressure is covered properly, including what actually moves it, on our high blood pressure page. Nothing here is a reason to change a prescription. A home reading is information for a conversation.
Body Shape and Heart Rate
Waist tells you the part the scales cannot. In 359,387 European adults followed a mean of 9.7 years, after adjusting for BMI, men in the highest fifth of waist circumference had 2.05 times the risk of death of those in the lowest, and women 1.78 times, while BMI stayed significant in the same models. Both are telling you something. Measure at the same landmark every time, standing, at the end of a normal breath out, tape snug without compressing. Across 120 studies the tape landmark made no substantial difference to what waist predicted at population level, so consistency matters for tracking yourself even though the site does not change the risk association.
Divide by your height in the same units: across 31 papers and more than 300,000 adults, waist-to-height ratio discriminated hypertension, diabetes and cardiovascular outcomes 4 to 5% better than BMI. A 35-inch (90 cm) waist on a 69-inch (175 cm) frame is 0.51. What waist means metabolically, through visceral fat and insulin resistance, is on insulin and glucose.
Resting heart rate is free and takes a minute: sit still for five minutes, then count your pulse for a full 60 seconds, at the same time of day, ideally before rising. Coffee, alcohol, illness, a poor night and a hot room all move it, so a single reading reflects the day more than your baseline. Pooled across 46 studies and 1,246,203 people, each 10 beats per minute higher resting heart rate came with a 9% higher rate of death from any cause, rising in a straight line from 45 beats per minute upward, so the risk inside the familiar 60 to 100 range is not flat.
Whether lowering it helps is the sharper question. In 29,325 Norwegian adults measured twice about a decade apart, a rise from under 70 to over 85 beats per minute carried an adjusted hazard ratio of 1.9 for ischaemic heart disease death, while a decrease over the same period showed no general mortality benefit. Beta blockers set the number for you, and an irregular rhythm makes both manual counting and cuffs unreliable. Heart rate variability is a different measure whose between-person spread makes comparing your figure to anyone else's close to meaningless, and we cover it on breathwork and HRV.
Sleep and Wearables
Sleep is the most measured thing in consumer health and the worst measured. Your own estimate of how long you sleep is a poor instrument: comparing self-report against wrist actigraphy in 669 middle-aged adults, measured sleep averaged 6.0 hours against 6.8 hours reported, the two correlated at only 0.45, and people sleeping five hours over-reported by about 1.3.
The measure with the better outcome data is sleep regularity, the day-to-day consistency of when you sleep and wake. Using more than 10 million hours of accelerometer data from 60,977 UK Biobank participants, the most regular four fifths of the sample had 20% to 48% lower all-cause mortality than the least regular fifth, and regularity predicted mortality more strongly than duration did. You can measure it without a device: write down the time you go to bed and the time you get up, every day for two weeks, on paper. The variation between days is the thing being measured. What paper cannot capture is the night you woke at three and lay there until five, the part that separates regularity from sleep quality, and sleep regularity covers how to steady it.
For wearables, steps are reliable: across 158 publications, Fitbit, Apple and Samsung devices counted steps accurately in laboratory conditions. Energy expenditure is not, with no brand accurate for calories burned, so treat that number as a motivational display. Sleep staging is where the marketing claims most exceed the instrument. Tested against polysomnography over three nights in 34 healthy young adults, consumer devices detected sleep well, with sensitivity of 0.93 or better, but caught wake poorly, with specificity between 0.18 and 0.54, and every device did worse on the disrupted nights.
A device bad at noticing you are awake will tell you that you slept on the night you know you did not. Readiness and recovery scores are proprietary arithmetic on top of those inputs, resting on the two things the hardware is worst at, with no outcome data of the kind behind the walking-speed or blood pressure numbers here. Total sleep time and step count are worth reading; the readiness score built on top of them is not.
There is a measurable cost. Among 172 patients with atrial fibrillation, the 83 who used a wearable reported more symptom monitoring and preoccupation, and 20% of them felt anxiety and always contacted a doctor in response to an irregular-rhythm notification, with significantly more ECGs, echocardiograms and ablations. It was retrospective and propensity-matched, so which way the anxiety runs is open, so weigh it before you strap on a monitor.
The Limits
Some tests are worth skipping, and some cannot be done at home at all.
The settled case for what not to measure is the general health check. Across 17 randomized trials and 251,891 participants, offering people a package of screening for many diseases and risk factors made no difference to total mortality (risk ratio 1.00, 95% CI 0.97 to 1.03) or to cancer mortality, on high-certainty evidence. Screening more things found more things, and finding them did not help. A specific test asked for a specific reason is a different thing entirely, and that conversation belongs with a doctor who knows your history.
The measure with the widest spread of all needs a lab. Cardiorespiratory fitness, measured on a treadmill with a technician, sorted mortality more sharply than almost anything: in 122,007 adults referred for exercise treadmill testing, the least fit quarter carried 5.04 times the risk-adjusted mortality of elite performers, larger than coronary artery disease, smoking or diabetes in the same model, with no upper limit of benefit found. That was a clinical referral population, so it reads as a general-population number only with care, and none of it can be reproduced with anything you own. Resting heart rate and self-reported walking pace are proxies for fitness and should be treated as such.
The reliability of the two Chinese medicine observations belongs here as well, stated plainly and held with respect. Thirty practitioners rating the same ten tongue slides reached the 80% agreement mark on 17.3% and 19.1% of occasions in two sessions, and only 5% of the time where the response choices were more complex; a later study using a formal classification scheme found agreement moderate for tongue coating and fair for tongue body color. Both studies trace the gap to the definitions rather than the observation, and agreement rose once practitioners used an explicit scheme. Pulse diagnosis is not a self-measure, and its reliability literature is limited even among practitioners: a review of twelve studies found acceptable agreement only where the method carried concrete operational definitions. These are findings about standardization, not verdicts on the art, and they sit in full below.
The Research & Studies
Everything here is based on the research we have collected and checked, sorted into groups and ordered with the strongest evidence first. Click any claim to open the studies behind it.
Progress Markers
Each 11 lb (5 kg) less grip strength came with a 16% higher death rate
Each 11 lb (5 kg) lower grip strength came with a 16% higher rate of death (HR 1.16, 95% CI 1.13 to 1.20) across 139,691 people in 17 countries.
Each 11 lb (5 kg) lower grip strength came with a 16% higher rate of death (HR 1.16, 95% CI 1.13 to 1.20) across 139,691 people in 17 countries. Measured in: 139,691 adults aged 35 to 70 across 17 countries, 58% women (the PURE study). What could explain it instead: Observational. Low grip strength is often an early sign of illness rather than its cause, which the short four-year follow-up makes more likely rather than less.. Grip is a marker of overall condition rather than a target in itself, so training your hands will not move the outcome. The often-repeated line that grip beats blood pressure as a predictor is narrower than it sounds: it was a post-hoc comparison, measured per standard deviation, and it holds for death but not for new cardiovascular disease, where blood pressure predicted better. Median follow-up was only four years, and these were adults from 35 upward rather than an older-adult sample.
Who this may not transfer to:Unusually well balanced for this literature, at 58% women, and the association held equally in both sexes.
The study · 1
Leong et al., prognostic value of grip strength, findings from the PURE study · Lancet 2015;386(9990):266-273
Counts once: this finding and 1 other here come from the same source, so they are one body of evidence, not separate confirmations.
Every 0.1 m/s faster usual walking speed came with 12% lower mortality
Every 0.1 m/s faster usual walking speed came with about 12% lower mortality (HR 0.88, 95% CI 0.87 to 0.90).
Every 0.1 m/s faster usual walking speed came with about 12% lower mortality (HR 0.88, 95% CI 0.87 to 0.90). Measured in: 34,485 community-dwelling adults aged 65 and over, pooled from nine cohorts, about 60% women. What could explain it instead: Observational, and slow walking reflects existing disease as much as it forecasts new disease.. A pooled analysis of individual participant data from nine selected cohorts rather than a systematic review, so it carries none of the publication-bias checks that phrase usually implies. Samples were mostly white and US-based. Predicted ten-year survival at 75 ran from 19% to 87% in men and 35% to 91% in women across the speed range, a remarkable spread for a test that takes seconds.
Who this may not transfer to:Results were reported separately for men and women, which is unusual and useful. Samples were predominantly white and US-based.
The study · 1
Studenski et al., gait speed and survival in older adults · JAMA 2011;305(1):50-58
Counts once: this finding and 1 other here come from the same source, so they are one body of evidence, not separate confirmations.
A regular cuff read 19.5 mmHg higher systolic on an arm needing an extra-large one
Against a correctly sized cuff, a regular cuff read 19.5 mmHg higher systolic (95% CI 16.1 to 22.9) in people whose arm required an extra-large cuff, 4.8 mmHg higher (3.0 to 6.6) in those requiring a large cuff, and 3.6 mmHg lower (-5.6 to -1.7) in those requiring a small cuff.
Against a correctly sized cuff, a regular cuff read 19.5 mmHg higher systolic (95% CI 16.1 to 22.9) in people whose arm required an extra-large cuff, 4.8 mmHg higher (3.0 to 6.6) in those requiring a large cuff, and 3.6 mmHg lower (-5.6 to -1.7) in those requiring a small cuff. Measured in: 195 community-dwelling adults in Baltimore across a range of mid-arm circumferences, mean age 54, 34% male, 68% Black, 51% with hypertension. Readings were taken with an automated oscillometric device in a research setting, so the figures isolate cuff error with every other variable controlled. At home the cuff error stacks with arm position and rest time rather than replacing them. Single-site and heavily Black, and mid-arm circumference distribution decides who is affected.
Who this may not transfer to:67 of 195 participants (34%) were male, so the sample is women-weighted; the abstract does not report the cuff-size error separately by sex.
The study · 1
Ishigami et al., effects of cuff size on the accuracy of blood pressure readings: the Cuff(SZ) randomized crossover trial · JAMA Intern Med 2023;183(10):1061-1068
After adjusting for BMI, the highest fifth of waist carried 2.05 times the death rate in men
After adjustment for BMI, men in the highest fifth of waist circumference had 2.05 times the risk of death of those in the lowest (95% CI 1.80 to 2.33) and women 1.78 times (1.56 to 2.04). For waist-to-hip ratio the figures were 1.68 and 1.51. BMI remained significantly associated with death in models that also contained waist.
After adjustment for BMI, men in the highest fifth of waist circumference had 2.05 times the risk of death of those in the lowest (95% CI 1.80 to 2.33) and women 1.78 times (1.56 to 2.04). For waist-to-hip ratio the figures were 1.68 and 1.51. BMI remained significantly associated with death in models that also contained waist. Measured in: 359,387 participants from nine European countries in the EPIC cohort, mean follow-up 9.7 years, 14,723 deaths; lowest mortality risk sat at a BMI of 25.3 in men and 24.3 in women. What could explain it instead: Reverse causation through illness-related weight loss. Undiagnosed cancer, COPD and heart failure shrink the waist years before they kill, which loads the leanest group with sick people and steepens the apparent gradient.. Waist and BMI are each adding information the other misses, so this is a case for measuring both rather than replacing one. The models adjust for education, smoking, alcohol, physical activity and height, and cannot adjust away the underlying problem that illness causes weight loss before it causes death.
Who this may not transfer to:Relative risks are reported separately for men and women throughout, and the BMI associated with lowest risk differed between them. The cohort is European, so transfer to populations with different body composition at the same waist is not established here.
The study · 1
Pischon et al., general and abdominal adiposity and risk of death in Europe · N Engl J Med 2008;359(20):2105-2120
The least fit fifth on a treadmill carried 5.04 times the mortality of elite performers
Risk-adjusted all-cause mortality fell across every fitness band with no upper limit found. The lowest fitness group carried 5.04 times the mortality of elite performers (95% CI 4.10 to 6.20), and below-average against above-average carried 1.41 (1.34 to 1.49). Those figures are comparable to or larger than coronary artery disease (1.29), smoking (1.41) and diabetes (1.40) in the same model.
Risk-adjusted all-cause mortality fell across every fitness band with no upper limit found. The lowest fitness group carried 5.04 times the mortality of elite performers (95% CI 4.10 to 6.20), and below-average against above-average carried 1.41 (1.34 to 1.49). Those figures are comparable to or larger than coronary artery disease (1.29), smoking (1.41) and diabetes (1.40) in the same model. Measured in: 122,007 adults referred for symptom-limited exercise treadmill testing at a US tertiary center between 1991 and 2014, mean age 53.4, 59.2% male, median follow-up 8.4 years, 13,637 deaths. What could explain it instead: Reverse causation. Undiagnosed illness lowers treadmill performance before it becomes a diagnosis, and the least fit group in a referral population is enriched for people who were already unwell on the day they were tested.. Everyone in this cohort was referred for a treadmill test, which means everyone had a clinical reason to be there, so it is not a general-population sample. Fitness was estimated from peak metabolic equivalents on a treadmill rather than measured by gas exchange, and none of it can be reproduced with anything you own.
Who this may not transfer to:72,173 of 122,007 (59.2%) were male, and fitness percentiles were age- and sex-matched by design, so the comparison groups are internally sex-adjusted.
The study · 1
Mandsager et al., association of cardiorespiratory fitness with long-term mortality among adults undergoing exercise treadmill testing · JAMA Netw Open 2018;1(6):e183605
People who could not hold a 10-second one-legged stance had 1.84 times the death rate
One in five could not hold the 10-second stance. Among them 17.5% died over the follow-up, against 4.6% of those who could, an adjusted hazard ratio of 1.84 (95% CI 1.23 to 2.78).
One in five could not hold the 10-second stance. Among them 17.5% died over the follow-up, against 4.6% of those who could, an adjusted hazard ratio of 1.84 (95% CI 1.23 to 2.78). Measured in: 1,702 adults aged 51 to 75 attending a private exercise-medicine clinic in Brazil, 68% men. What could explain it instead: Observational. Neurological and musculoskeletal disease impairs balance and independently raises mortality, and a clinic population differs from the general public in ways adjustment cannot remove.. Much the smallest of our three self-measures, and not a general-population sample: these were self-selected attendees at one clinic. The confidence interval is wide. Treat it as a free signal worth acting on rather than a precise multiplier, and not as equal in weight to the grip and walking-speed findings.
Who this may not transfer to:Just over two thirds of this sample were men. Balance, falls and the muscle loss underneath them differ between men and women, and women carry higher fracture risk, so the size of this association in women is not established here.
The study · 1
Araujo et al., successful 10-second one-legged stance performance predicts survival · Br J Sports Med 2022;56(17):975-980
Walking speed that improved over a year meant 31.6% eight-year mortality against 49.3%
Of six health and function measures tracked quarterly for a year, only improved gait speed predicted eight-year survival. Mortality was 31.6% in those who improved by at least 0.1 m/s, 41.2% in those who improved transiently, and 49.3% in those who never improved (adjusted hazard ratio 0.42, 95% CI 0.29 to 0.61).
Of six health and function measures tracked quarterly for a year, only improved gait speed predicted eight-year survival. Mortality was 31.6% in those who improved by at least 0.1 m/s, 41.2% in those who improved transiently, and 49.3% in those who never improved (adjusted hazard ratio 0.42, 95% CI 0.29 to 0.61). Measured in: 439 adults aged 65 and over recruited through a Medicare health maintenance organization and Veterans Affairs primary care programs in the United States. What could explain it instead: Observational. Recovery from an acute illness or a hospitalization raises gait speed and independently improves survival, so improvement partly marks who was temporarily unwell at baseline rather than who got fitter.. This shows that people whose walking speed improves live longer, not that making your walking speed improve makes you live longer. The authors end by calling for research into whether interventions that raise gait speed affect survival, which had not been done. The sample is small, single-country, and drawn from two specific health systems.
Who this may not transfer to:The abstract does not give the sex split, and reports only that the survival benefit was consistent across subgroups based on age, sex, ethnicity, initial gait speed, healthcare system and hospitalization.
The study · 1
Hardy et al., improvement in usual gait speed predicts better survival in older adults · J Am Geriatr Soc 2007;55(11):1727-1734
Grip strength showed no association with fracture, fall injury or new diabetes
In the same cohort where each 11 lb (5 kg) less grip carried a 16% higher rate of death, grip strength showed no significant association with incident diabetes, with hospital admission for pneumonia or COPD, with injury due to a fall, or with fracture.
In the same cohort where each 11 lb (5 kg) less grip carried a 16% higher rate of death, grip strength showed no significant association with incident diabetes, with hospital admission for pneumonia or COPD, with injury due to a fall, or with fracture. Measured in: 139,691 adults aged 35 to 70 across 17 countries, 58% women, grip measured with a Jamar dynamometer, median follow-up 4.0 years. What could explain it instead: Observational. Fracture and fall injury depend heavily on bone density, medication and home environment, none of which grip strength measures, so a null here reflects what the test does not read rather than an absence of relationship in the body.. A null over a median four years of follow-up is a weaker statement than a null over fifteen, particularly for outcomes as uncommon as fracture in a cohort starting at 35. The finding is that grip is a general marker rather than a specific one, and it does not license the reverse claim that hand strength is unrelated to falls.
Who this may not transfer to:Unusually well balanced for this literature at 58% women, across 17 countries of varying income.
The study · 1
Leong et al., prognostic value of grip strength: findings from the PURE study · Lancet 2015;386(9990):266-273
Counts once: this finding and 1 other here come from the same source, so they are one body of evidence, not separate confirmations.
With both in one model only home blood pressure still predicted events (HR 1.22), not office
Each entered alone, office pressure (hazard ratio 1.13 per 10 mmHg systolic) and home pressure (1.23) both predicted cardiovascular events. Entered into the same model, only home pressure still predicted events (1.22, 95% CI 1.09 to 1.37) and office pressure did not (1.01, 0.92 to 1.12). Systolic home pressure was the sole predictor of total mortality.
Each entered alone, office pressure (hazard ratio 1.13 per 10 mmHg systolic) and home pressure (1.23) both predicted cardiovascular events. Entered into the same model, only home pressure still predicted events (1.22, 95% CI 1.09 to 1.37) and office pressure did not (1.01, 0.92 to 1.12). Systolic home pressure was the sole predictor of total mortality. Measured in: 2,081 randomly selected Finnish adults aged 45 to 74, followed a mean of 6.8 years, with 162 cardiovascular events and 118 deaths. What could explain it instead: Observational. People willing and able to self-measure at home differ in health literacy, medication adherence and comorbidity from those who are not, and that difference tracks cardiovascular risk independently of the reading.. 162 events is a modest number for a model comparison, and the confidence intervals overlap more than a single point estimate suggests. Home readings in this study were taken under instruction with a supplied validated device, which is not the same as readings taken on whatever monitor a person owns. Ambulatory monitoring, not home monitoring, remains the reference standard for out-of-office pressure.
Who this may not transfer to:The abstract does not report the sex composition of the randomly selected national sample or give sex-stratified hazard ratios.
The studies · 2
Niiranen et al., home-measured blood pressure is a stronger predictor of cardiovascular risk than office blood pressure: the Finn-Home study · Hypertension 2010;55(6):1346-1351
Muntner et al., measurement of blood pressure in humans: a scientific statement from the American Heart Association · Hypertension 2019;73(5):e35-e66
Each 10 beats per minute higher resting heart rate carried a 9% higher death rate
Each 10 beats per minute higher resting heart rate carried a relative risk of 1.09 (95% CI 1.07 to 1.12) for all-cause mortality and 1.08 (1.06 to 1.10) for cardiovascular mortality. Against the lowest category, 60 to 80 beats per minute carried 1.12 and above 80 beats per minute carried 1.45 (1.34 to 1.57). All-cause mortality risk rose linearly from 45 beats per minute upward.
Each 10 beats per minute higher resting heart rate carried a relative risk of 1.09 (95% CI 1.07 to 1.12) for all-cause mortality and 1.08 (1.06 to 1.10) for cardiovascular mortality. Against the lowest category, 60 to 80 beats per minute carried 1.12 and above 80 beats per minute carried 1.45 (1.34 to 1.57). All-cause mortality risk rose linearly from 45 beats per minute upward. Measured in: 46 prospective cohort studies in general populations, 1,246,203 participants and 78,349 deaths for all-cause mortality; 848,320 participants and 25,800 deaths for cardiovascular mortality. What could explain it instead: Observational. Fitness, thyroid disease, anemia, fever, subclinical heart failure and beta blocker use all set resting heart rate and independently affect mortality, and adjustment for traditional cardiovascular risk factors does not touch most of them.. The authors report substantial heterogeneity between studies and detected publication bias, which is why this sits at moderate rather than strong despite the size. Resting heart rate is also measured inconsistently across cohorts, and the cardiovascular risk increase only reached significance at 90 beats per minute.
Who this may not transfer to:The pooled analysis does not report a sex breakdown across its 46 cohorts or give sex-stratified dose-response curves.
The study · 1
Zhang, Shen and Qi, resting heart rate and all-cause and cardiovascular mortality in the general population: a meta-analysis · CMAJ 2016;188(3):E53-E63
A resting heart rate rising from under 70 to over 85 over a decade carried a 1.9 hazard ratio
Against people under 70 beats per minute at both measurements (8.2 ischaemic heart disease deaths per 10,000 person-years), those under 70 at first measurement and above 85 ten years later had an adjusted hazard ratio of 1.9 (95% CI 1.0 to 3.6), and those going from 70 to 85 up to above 85 had 1.8 (1.2 to 2.8). The association was not linear (p = 0.003 for quadratic trend), and a decrease in resting heart rate showed no general mortality benefit.
Against people under 70 beats per minute at both measurements (8.2 ischaemic heart disease deaths per 10,000 person-years), those under 70 at first measurement and above 85 ten years later had an adjusted hazard ratio of 1.9 (95% CI 1.0 to 3.6), and those going from 70 to 85 up to above 85 had 1.8 (1.2 to 2.8). The association was not linear (p = 0.003 for quadratic trend), and a decrease in resting heart rate showed no general mortality benefit. Measured in: 13,499 men and 15,826 women in Norway without known cardiovascular disease, measured twice about ten years apart in the Nord-Trondelag health study, mean 12 years of follow-up, 3,038 deaths of which 388 from ischaemic heart disease. What could explain it instead: Observational. A resting heart rate that climbs over a decade is a signal of developing illness, weight gain, deconditioning or new medication, all of which raise cardiac mortality on their own, so the rise marks the process rather than causing it.. 388 ischaemic heart disease deaths across many change categories makes the individual hazard ratios imprecise, and the confidence interval on the largest of them touches 1.0. The finding that a falling rate carried no general benefit is the part most relevant to anyone trying to train the number down, and it is a null within an observational study rather than a trial of lowering it.
Who this may not transfer to:Both sexes were enrolled in near-equal numbers, 13,499 men and 15,826 women, and the conclusion is stated for men and women together.
The study · 1
Nauman et al., temporal changes in resting heart rate and deaths from ischemic heart disease · JAMA 2011;306(23):2579-2587
The most regular four fifths of sleepers had 20% to 48% lower all-cause mortality
Across the top four quintiles of the Sleep Regularity Index against the least regular quintile, all-cause mortality was 20% to 48% lower, cancer mortality 16% to 39% lower and cardiometabolic mortality 22% to 57% lower. Sleep regularity predicted all-cause mortality more strongly than sleep duration did, on both model comparisons run.
Across the top four quintiles of the Sleep Regularity Index against the least regular quintile, all-cause mortality was 20% to 48% lower, cancer mortality 16% to 39% lower and cardiometabolic mortality 22% to 57% lower. Sleep regularity predicted all-cause mortality more strongly than sleep duration did, on both model comparisons run. Measured in: 60,977 UK Biobank participants, mean age 62.8, 55.0% female, with more than 10 million hours of accelerometer data and 1,859 deaths over a mean 6.3 years. What could explain it instead: Observational. Shift work, chronic pain, depression and terminal illness all fragment sleep timing and independently raise mortality, so irregularity partly marks who is unwell or working nights.. Six years of follow-up in a cohort with a median regularity index of 81 leaves little room to separate irregular sleep from the shift work, illness and social circumstances that produce it. The comparison with duration rests on nested model tests whose p values (0.14 to 0.20) show duration adding nothing, rather than showing regularity to be a cause. UK Biobank participants are healthier and less deprived than the UK population.
Who this may not transfer to:55.0% female, and results were adjusted for sex. The cohort skews older, white and British, so the size of the association in younger or shift-working populations is not established here.
The study · 1
Windred et al., sleep regularity is a stronger predictor of mortality risk than sleep duration: a prospective cohort study · Sleep 2024;47(1):zsad253
Scoring 0 to 3 on the sitting-rising test carried 5.44 times the death rate of a top score
Scored 0 to 10 from how many hand or knee supports you need to lower yourself to the floor and stand back up, lower scores carried higher mortality. Against the top band of 8 to 10, multivariate-adjusted hazard ratios were 5.44 (95% CI 3.1 to 9.5) for scores of 0 to 3, 3.44 (2.0 to 5.9) for 3.5 to 5.5 and 1.84 (1.1 to 3.0) for 6 to 7.5. Each one-point increase carried a 21% improvement in survival.
Scored 0 to 10 from how many hand or knee supports you need to lower yourself to the floor and stand back up, lower scores carried higher mortality. Against the top band of 8 to 10, multivariate-adjusted hazard ratios were 5.44 (95% CI 3.1 to 9.5) for scores of 0 to 3, 3.44 (2.0 to 5.9) for 3.5 to 5.5 and 1.84 (1.1 to 3.0) for 6 to 7.5. Each one-point increase carried a 21% improvement in survival. Measured in: 2,002 adults aged 51 to 80, 68% men, attending the CLINIMEX exercise-medicine clinic in Rio de Janeiro, median follow-up 6.3 years, 159 deaths (7.9%). What could explain it instead: Observational. The conditions that make it hard to get off the floor, arthritis, obesity, and neurological and cardiac disease among them, independently raise mortality, so the score reads existing illness as much as it forecasts new illness.. The same self-selected clinic population as the ten-second balance finding, so it is not a general-population sample and the confidence intervals are wide. The score combines strength, flexibility and body weight, so a low result does not point to any one cause, and the test asks you to lower yourself to the floor, which is not safe to attempt alone if you are frail or have fallen in the past year.
Who this may not transfer to:68% of this clinic sample were men, and the study does not report the mortality gradient separately by sex. It is a Brazilian exercise-medicine referral population rather than a general one, so the size of the association elsewhere is not established here.
The study · 1
Brito et al., ability to sit and rise from the floor as a predictor of all-cause mortality · Eur J Prev Cardiol 2014;21(7):892-898
Measurement And Diagnosis
Walking speed is measured over a 13-foot (4-metre) course, where the pooled average was 0.92 m/s
The pooled survival analysis was built on a stopwatch and a 13-foot (4-metre) course: from a standing start, walk at your usual pace, as if walking down the street, with no further encouragement. Mean speed in the 34,485 participants was 0.92 m/s (SD 0.27).
The pooled survival analysis was built on a stopwatch and a 13-foot (4-metre) course: from a standing start, walk at your usual pace, as if walking down the street, with no further encouragement. Mean speed in the 34,485 participants was 0.92 m/s (SD 0.27). Measured in: 34,485 community-dwelling adults aged 65 and over, mean age 73.5, 59.6% women, 79.8% white, pooled from nine cohorts collected between 1986 and 2000. What could explain it instead: Observational. Slow walking is a common final pathway for heart failure, arthritis, neuropathy and depression, so the speed reads existing disease as much as it forecasts new disease.. The contributing cohorts did not all use a 13-foot (4-metre) course. Walk distances ran from 8 feet to 20 feet (6 metres) and were converted to a 13-foot (4-metre) equivalent by formula, including one conversion derived from just 61 people walking both distances. Reproducing the protocol at home gets you closer to the published figures than any variation of it, and the pooled number still carries that conversion inside it.
Who this may not transfer to:Results were reported separately by sex, which is unusual and useful: predicted ten-year survival at 75 ranged from 19% to 87% in men and 35% to 91% in women across the speed range. Samples were predominantly white and US-based.
The study · 1
Studenski et al., gait speed and survival in older adults · JAMA 2011;305(1):50-58
Counts once: this finding and 1 other here come from the same source, so they are one body of evidence, not separate confirmations.
An unsupported arm read 6.5 mmHg higher systolic than one supported at heart height
Against the arm supported on a desk at heart height, resting the hand on the lap overestimated systolic pressure by 3.9 mmHg (95% CI 2.5 to 5.2) and diastolic by 4.0 mmHg (3.1 to 5.0). An unsupported arm at the side overestimated systolic by 6.5 mmHg (5.1 to 7.9) and diastolic by 4.4 mmHg (3.4 to 5.4).
Against the arm supported on a desk at heart height, resting the hand on the lap overestimated systolic pressure by 3.9 mmHg (95% CI 2.5 to 5.2) and diastolic by 4.0 mmHg (3.1 to 5.0). An unsupported arm at the side overestimated systolic by 6.5 mmHg (5.1 to 7.9) and diastolic by 4.4 mmHg (3.4 to 5.4). Measured in: 133 adults aged 18 to 80 recruited in Baltimore, mean age 57, 53% female; 36% had systolic pressure of 130 mmHg or above and 41% had a BMI of 30 or above. Single-site trial with triplicate automated readings under research conditions, so it isolates arm position with everything else held constant. It does not say how often each position is used in practice, which is the quantity that decides the population-level error.
Who this may not transfer to:70 of 133 participants (53%) were female. The trial reports results as consistent across demographic subgroups.
The study · 1
Liu et al., arm position and blood pressure readings: the ARMS crossover randomized clinical trial · JAMA Intern Med 2024;184(12):1436-1442
General health checks made no difference to total mortality (risk ratio 1.00)
General health checks had little or no effect on total mortality (risk ratio 1.00, 95% CI 0.97 to 1.03; 11 trials, 233,298 participants, 21,535 deaths, high certainty, I-squared 0%) or cancer mortality (1.01, 0.92 to 1.12, high certainty), and probably little or none on cardiovascular mortality (1.05, 0.94 to 1.16, moderate certainty). Ischaemic heart disease events were unchanged (0.98, 0.94 to 1.03, high certainty).
General health checks had little or no effect on total mortality (risk ratio 1.00, 95% CI 0.97 to 1.03; 11 trials, 233,298 participants, 21,535 deaths, high certainty, I-squared 0%) or cancer mortality (1.01, 0.92 to 1.12, high certainty), and probably little or none on cardiovascular mortality (1.05, 0.94 to 1.16, moderate certainty). Ischaemic heart disease events were unchanged (0.98, 0.94 to 1.03, high certainty). Measured in: 17 randomized trials, 15 reporting outcome data, 251,891 participants unselected for disease or risk factors; geriatric trials were excluded by design. This is about packaged screening of more than one organ system in people with no particular reason to be tested. It says nothing about a specific test ordered for a specific reason, nothing about established single-disease screening programmes, and nothing about older adults, who were excluded. Several included trials are old enough that the treatments available on a positive result have since improved.
Who this may not transfer to:The review does not report a pooled sex breakdown or sex-stratified mortality effects across its 17 trials.
The study · 1
Krogsboll, Jorgensen and Gotzsche, general health checks in adults for reducing morbidity and mortality from disease · Cochrane Database Syst Rev 2019;1(1):CD009009
The 30-second chair stand tracked leg-press strength at r = 0.71 to 0.78
For the 30-second chair stand, test-retest intraclass correlation was 0.84 in men and 0.92 in women, and chair-stand count correlated with maximum weight-adjusted leg press at r = 0.78 in men and r = 0.71 in women. Performance fell significantly across each decade from the 60s to the 80s and was lower in low-active than high-active participants.
For the 30-second chair stand, test-retest intraclass correlation was 0.84 in men and 0.92 in women, and chair-stand count correlated with maximum weight-adjusted leg press at r = 0.78 in men and r = 0.71 in women. Performance fell significantly across each decade from the 60s to the 80s and was lower in low-active than high-active participants. Measured in: 76 community-dwelling volunteers over 60, mean age 70.5, for the reliability and validity work; the criterion-referenced independence standards were later derived from 2,140 moderate-functioning older adults aged 60 to 94. What could explain it instead: Volunteer selection. People who agree to perform two maximal leg-press tests are fitter and more mobile than the population the test is used to screen, which inflates both the reliability and the correlation with leg press.. The validation sample was 76 generally active volunteers, and the authors limit their conclusion to generally active community-dwelling older adults. The published cut-offs describe the capacity associated with staying physically independent, which is a different question from whether you will fall, and neither paper measured falls.
Who this may not transfer to:Reliability and validity were reported separately for men and women in the original validation, and the independence standards are published separately by sex for ages 60 to 94.
The studies · 2
Jones, Rikli and Beam, a 30-s chair-stand test as a measure of lower body strength in community-residing older adults · Res Q Exerc Sport 1999;70(2):113-119
Rikli and Jones, development and validation of criterion-referenced clinically relevant fitness standards for maintaining physical independence in later years · Gerontologist 2013;53(2):255-267
Waist-to-height ratio discriminated cardiometabolic risk 4 to 5% better than BMI
Pooling receiver operating characteristic data, waist circumference improved discrimination of adverse cardiometabolic outcomes by 3% over BMI and waist-to-height ratio by 4 to 5% over BMI. Within-study comparison of the area under the curve found waist-to-height ratio significantly better than waist circumference for diabetes, hypertension, cardiovascular disease and all outcomes combined, in both sexes.
Pooling receiver operating characteristic data, waist circumference improved discrimination of adverse cardiometabolic outcomes by 3% over BMI and waist-to-height ratio by 4 to 5% over BMI. Within-study comparison of the area under the curve found waist-to-height ratio significantly better than waist circumference for diabetes, hypertension, cardiovascular disease and all outcomes combined, in both sexes. Measured in: 31 papers covering more than 300,000 adults across several ethnic groups. These are discrimination statistics for detecting risk factors that are already present, not prospective prediction of events, and a 4 to 5% improvement in area under the curve is a modest gain over a measure most people already have. The included studies used different waist landmarks and different outcome definitions.
Who this may not transfer to:The superiority of waist-to-height ratio is reported as holding in both men and women, across several nationalities and ethnic groups.
The study · 1
Ashwell, Gunn and Gibson, waist-to-height ratio is a better screening tool than waist circumference and BMI for adult cardiometabolic risk factors: systematic review and meta-analysis · Obes Rev 2012;13(3):275-286
Across 120 studies the tape landmark made no substantial difference to what waist predicted
Across 120 studies and 236 samples, the measurement protocol had no substantial influence on the association between waist circumference and all-cause or cardiovascular mortality, cardiovascular disease or diabetes. Significant associations were found in 65% of samples, and the common protocols (minimal waist, midpoint between rib and hip, umbilicus) were distributed similarly among the significant and non-significant ones.
Across 120 studies and 236 samples, the measurement protocol had no substantial influence on the association between waist circumference and all-cause or cardiovascular mortality, cardiovascular disease or diabetes. Significant associations were found in 65% of samples, and the common protocols (minimal waist, midpoint between rib and hip, umbilicus) were distributed similarly among the significant and non-significant ones. Measured in: 120 studies, 236 samples, analyzed by an expert panel across sample size, sex, age, race and ethnicity. This says the site does not change the risk association at population level. It does not say the sites give the same number: they do not, and switching landmarks between your own measurements will produce a change that is entirely artefact. The review also notes that no scientific rationale has been published for any of the protocols recommended by major health authorities.
Who this may not transfer to:Similar patterns of association were observed across sample size, sex, age, race and ethnicity, which is the specific comparison the panel set out to make.
The study · 1
Ross et al., does the relationship between waist circumference, morbidity and mortality depend on measurement protocol for waist circumference? · Obes Rev 2008;9(4):312-325
Reported sleep matched measured sleep at only r = 0.45, over-reporting by about 0.8 hours
Average measured sleep was 6.0 hours against 6.8 hours reported, a gap of about 0.8 hours. Reports rose by only 31 minutes for each additional hour of measured sleep, and the correlation between reported and measured duration was 0.45. Modelled, people sleeping 5 hours over-reported by 1.3 hours and those sleeping 7 hours over-reported by 0.3 hours.
Average measured sleep was 6.0 hours against 6.8 hours reported, a gap of about 0.8 hours. Reports rose by only 31 minutes for each additional hour of measured sleep, and the correlation between reported and measured duration was 0.45. Modelled, people sleeping 5 hours over-reported by 1.3 hours and those sleeping 7 hours over-reported by 0.3 hours. Measured in: 669 middle-aged adults at the Chicago site of the CARDIA study, 82% of those invited, measured with three days each of wrist actigraphy, a sleep log and questions about usual sleep duration in two waves. What could explain it instead: Health, sociodemographic and sleep characteristics all altered the size of the discrepancy in this sample, so the over-reporting is not a constant offset that can be subtracted, and groups differing in health differ in how wrong their estimate is.. Actigraphy is itself an estimate: it infers sleep from movement and tends to score quiet wakefulness as sleep, so the true gap between belief and sleep may differ from 0.8 hours in either direction. One city, one age band, and three nights per wave.
Who this may not transfer to:The abstract does not give the sex split, and reports only that the size of the discrepancy varied by health, sociodemographic and sleep characteristics.
The study · 1
Lauderdale et al., self-reported and measured sleep duration: how similar are they? · Epidemiology 2008;19(6):838-845
Sleep trackers detected sleep well but caught wake poorly, at 0.18 to 0.54 specificity
Epoch-by-epoch sensitivity for detecting sleep was high across all devices (0.93 or better) while specificity for detecting wake was low to medium (0.18 to 0.54). Sleep stage comparisons were mixed, and devices performed worse on the disrupted-sleep night. Most devices matched or beat research actigraphy on sleep and wake; two Garmin devices performed worse.
Epoch-by-epoch sensitivity for detecting sleep was high across all devices (0.93 or better) while specificity for detecting wake was low to medium (0.18 to 0.54). Sleep stage comparisons were mixed, and devices performed worse on the disrupted-sleep night. Most devices matched or beat research actigraphy on sleep and wake; two Garmin devices performed worse. Measured in: 34 healthy young adults, 22 women, mean age 28.1, tested on three consecutive nights in a sleep laboratory including one disrupted-sleep condition, against polysomnography. Healthy young adults sleeping in a laboratory is close to the easiest condition a sleep tracker will ever face, and the devices already failed at detecting wake. Anyone with insomnia, apnea or fragmented sleep is outside what was tested. Consumer firmware is also revised continuously, so a device tested in 2021 is not the device sold now.
Who this may not transfer to:22 of 34 participants were women, which is the reverse of the usual skew, and the study does not report device accuracy separately by sex.
The study · 1
Chinoy et al., performance of seven consumer sleep-tracking devices compared with polysomnography · Sleep 2021;44(5):zsaa291
Wearables counted steps accurately, but no brand measured energy expenditure accurately
Across 158 publications covering nine commercial device brands, Fitbit, Apple and Samsung devices counted steps accurately in controlled laboratory conditions. Heart rate accuracy varied by brand, with Apple Watch and Garmin most accurate and Fitbit tending to underestimate. No brand was accurate for energy expenditure.
Across 158 publications covering nine commercial device brands, Fitbit, Apple and Samsung devices counted steps accurately in controlled laboratory conditions. Heart rate accuracy varied by brand, with Apple Watch and Garmin most accurate and Fitbit tending to underestimate. No brand was accurate for energy expenditure. Measured in: 158 publications examining nine commercial wearable device brands, in laboratory and free-living conditions. Accuracy figures come mostly from controlled laboratory testing, which flatters step counting: slow walking, pushing a trolley and wheelchair use are the conditions where counters fail, and they are underrepresented. Devices are redesigned faster than they are validated, so a brand-level verdict ages quickly.
Who this may not transfer to:The review does not pool a sex breakdown across its included studies or report device accuracy separately for women, which matters because wrist size and skin tone both affect optical heart rate sensing.
The study · 1
Fuller et al., reliability and validity of commercially available wearable devices for measuring steps, energy expenditure, and heart rate: systematic review · JMIR Mhealth Uhealth 2020;8(9):e18694
Practitioners reached 80% agreement on a tongue only 5% of the time for complex features
Thirty practitioners rating ten tongue slides reached the 80% agreement threshold on 17.3% of occasions in the first session and 19.1% in the second, and almost all of those were simple yes-or-no questions: with more complex response choices, 80% agreement was reached 5% of the time. In a later study using a formal classification scheme, inter-rater reliability across 17 clinicians was moderate for tongue coating (Gwet AC2 0.49 to 0.55) and fair for tongue body color and other body features (0.34).
Thirty practitioners rating ten tongue slides reached the 80% agreement threshold on 17.3% of occasions in the first session and 19.1% in the second, and almost all of those were simple yes-or-no questions: with more complex response choices, 80% agreement was reached 5% of the time. In a later study using a formal classification scheme, inter-rater reliability across 17 clinicians was moderate for tongue coating (Gwet AC2 0.49 to 0.55) and fair for tongue body color and other body features (0.34). Measured in: 30 traditional Chinese medicine practitioners rating 10 tongue slides across two sessions; separately, 17 clinicians in Hong Kong and mainland China rating 24 representative smartphone tongue images. What could explain it instead: Terminology rather than perception. Practitioners trained in different lineages and languages use different operational definitions and different tongue regions, so disagreement about a word is scored as disagreement about a tongue.. Both studies conclude that the problem is the definitions rather than the observation: agreement rose when practitioners used an explicit operating classification scheme, and the same-rater agreement between looking at a person and looking at the photograph was good to very good. This measures agreement about a description, and it does not test whether tongue signs carry clinical information.
Who this may not transfer to:Neither study reports the sex of the practitioners or of the people whose tongues were photographed. The unit of analysis is the rater rather than a patient population.
The studies · 2
Kim, Cobbin and Zaslawski, traditional Chinese medicine tongue inspection: an examination of the inter- and intrapractitioner reliability for specific tongue characteristics · J Altern Complement Med 2008;14(5):527-536
Wang et al., intra-rater and inter-rater reliability of tongue coating diagnosis in traditional Chinese medicine using smartphones: quasi-Delphi study · JMIR Mhealth Uhealth 2020;8(7):e16018
As a fall predictor the timed up and go had 0.31 sensitivity, missing two thirds of fallers
At the usual threshold of 13.5 seconds or more, pooled specificity was 0.74 (95% CI 0.52 to 0.88) and pooled sensitivity 0.31 (95% CI 0.13 to 0.57), so the test misses roughly two thirds of the people who go on to fall. The review's own logistic regression found the score was not a significant predictor of falls (OR 1.01, 95% CI 1.00 to 1.02, p = 0.05).
At the usual threshold of 13.5 seconds or more, pooled specificity was 0.74 (95% CI 0.52 to 0.88) and pooled sensitivity 0.31 (95% CI 0.13 to 0.57), so the test misses roughly two thirds of the people who go on to fall. The review's own logistic regression found the score was not a significant predictor of falls (OR 1.01, 95% CI 1.00 to 1.02, p = 0.05). Measured in: 25 studies reviewed in community-dwelling older adults, of which 10 could be pooled for the meta-analysis. This is about predicting falls in community-dwelling older adults specifically. The test remains a valid measure of mobility, correlating with the Berg Balance Scale and with activities-of-daily-living scores in its original validation, and the review's conclusion is that it should not be used in isolation to identify people at high risk of falling rather than that it measures nothing.
Who this may not transfer to:The review does not report a pooled sex breakdown across its 25 studies. The original 1991 description of the test used 60 geriatric day-hospital patients with a mean age of 79.5.
The studies · 2
Barry et al., is the Timed Up and Go test a useful predictor of risk of falls in community dwelling older adults: a systematic review and meta-analysis · BMC Geriatr 2014;14:14
Podsiadlo and Richardson, the timed Up and Go: a test of basic functional mobility for frail elderly persons · J Am Geriatr Soc 1991;39(2):142-148
One in five wearable users felt anxiety and always called a doctor over a rhythm alert
Wearable users reported higher rates of symptom monitoring and preoccupation (p = 0.03) and more treatment concerns (p = 0.02) than non-users. 20% of wearable users experienced anxiety and always contacted their doctor in response to an irregular rhythm notification. After matching, atrial-fibrillation-specific health care use was significantly greater in users (p = 0.04), including more ECGs, echocardiograms and ablations.
Wearable users reported higher rates of symptom monitoring and preoccupation (p = 0.03) and more treatment concerns (p = 0.02) than non-users. 20% of wearable users experienced anxiety and always contacted their doctor in response to an irregular rhythm notification. After matching, atrial-fibrillation-specific health care use was significantly greater in users (p = 0.04), including more ECGs, echocardiograms and ablations. Measured in: 172 patients with atrial fibrillation at a US academic center, mean age 72.6, 42% women, of whom 83 used a wearable, followed over 9 months of merged survey and electronic health record data. What could explain it instead: Selection by health anxiety. People already preoccupied with their rhythm are more likely to buy a monitor and more likely to call a doctor, which would produce this entire association without the device causing anything.. Retrospective and propensity-matched, so the direction of causation is open: more anxious patients may be more likely to buy a wearable in the first place. 172 people in a US academic health system, and the authors call for prospective randomized study before concluding anything about net effect. This is a cardiac population, and it says nothing directly about readiness scores in healthy users.
Who this may not transfer to:42% women. The sample is small enough that sex-stratified estimates are not reported, and anxiety about cardiac symptoms is documented as differing between men and women.
The study · 1
Rosman et al., wearable devices, health care use, and psychological well-being in patients with atrial fibrillation · J Am Heart Assoc 2024;13(15):e033750
Pulse-diagnosis agreement held only where the method carried concrete operational definitions
Twelve eligible studies were found: three evaluated both intra- and inter-rater reliability and nine only inter-rater. Acceptable agreement was achieved where the method carried concrete operational definitions. Poor agreement traced to unclear definitions and terminology in the classical descriptions and to imprecise descriptions persisting in standardized systems. Most studies did not assess intra-rater reliability at all.
Twelve eligible studies were found: three evaluated both intra- and inter-rater reliability and nine only inter-rater. Acceptable agreement was achieved where the method carried concrete operational definitions. Poor agreement traced to unclear definitions and terminology in the classical descriptions and to imprecise descriptions persisting in standardized systems. Most studies did not assess intra-rater reliability at all. Measured in: 12 studies of manual pulse diagnosis at the radial artery by human testers, published in English. A narrative review with no pooling, over a literature the authors describe as consistently limited by small samples and by testers who often knew the participants beforehand. It reports how far practitioners agree with each other, which is a prerequisite for a diagnostic method rather than evidence about what the method detects.
Who this may not transfer to:The review does not report the sex of testers or participants in the included studies.
The study · 1
Bilton and Zaslawski, reliability of manual pulse diagnosis methods in traditional East Asian medicine: a systematic narrative literature review · J Altern Complement Med 2016;22(8):599-609
Heart And Vascular
Self-monitoring used to titrate medication lowered systolic pressure by 3.5 mmHg
At twelve months, systolic pressure was 3.5 mmHg lower than usual care with self-monitoring alone (95% CI -5.8 to -1.2) and 4.7 mmHg lower with telemonitoring added (-7.0 to -2.4). Telemonitoring did not beat self-monitoring alone (-1.2 mmHg, 95% CI -3.5 to 1.2). Adverse events were similar across all three groups.
At twelve months, systolic pressure was 3.5 mmHg lower than usual care with self-monitoring alone (95% CI -5.8 to -1.2) and 4.7 mmHg lower with telemonitoring added (-7.0 to -2.4). Telemonitoring did not beat self-monitoring alone (-1.2 mmHg, 95% CI -3.5 to 1.2). Adverse events were similar across all three groups. Measured in: 1,182 patients over 35 with blood pressure above 140/90 mmHg, across 142 UK general practices, who were willing to self-monitor; 1,003 (85%) were included in the primary analysis. The effect came from GPs titrating medication on the readings, not from the readings existing, so self-monitoring with nobody adjusting anything is a different intervention with no reason to expect this result. Participants were selected for willingness to self-monitor and the trial was unmasked. The monitor manufacturer was among the funders.
Who this may not transfer to:The abstract does not report the sex split of the randomized sample or give sex-stratified blood pressure differences.
The study · 1
McManus et al., efficacy of self-monitored blood pressure, with or without telemonitoring, for titration of antihypertensive medication (TASMINH4) · Lancet 2018;391(10124):949-959
Muscle And Strength
Progressive resistance training raised gait speed by 0.08 m/s and improved chair rise
Progressive resistance training improved gait speed by 0.08 m/s (24 trials, 1,179 participants, 95% CI 0.04 to 0.12) and produced a moderate to large improvement in getting out of a chair (11 trials, 384 participants, SMD -0.94, 95% CI -1.49 to -0.38). Muscle strength improved substantially (73 trials, 3,059 participants, SMD 0.84). Physical ability improved slightly (33 trials, SMD 0.14).
Progressive resistance training improved gait speed by 0.08 m/s (24 trials, 1,179 participants, 95% CI 0.04 to 0.12) and produced a moderate to large improvement in getting out of a chair (11 trials, 384 participants, SMD -0.94, 95% CI -1.49 to -0.38). Muscle strength improved substantially (73 trials, 3,059 participants, SMD 0.84). Physical ability improved slightly (33 trials, SMD 0.14). Measured in: 121 randomized trials, 6,700 participants, mostly older adults, with training typically two to three times a week at high intensity. These are improvements in the test scores, not in the outcomes the test scores predict, and no trial here measured survival. Adverse events were poorly recorded across the literature, though joint pain and muscle soreness were commonly reported where they were tracked, and serious events were rare and not attributed to the exercise.
Who this may not transfer to:The review does not report a pooled sex breakdown across its 121 trials or give sex-stratified effect sizes for gait speed or chair rise.
The study · 1
Liu and Latham, progressive resistance strength training for improving physical function in older adults · Cochrane Database Syst Rev 2009;(3):CD002759
How to Use These
You can do the physical set in one morning. Before getting out of bed, count your pulse for a full minute. Sit for five minutes, then take blood pressure with the correct cuff and the arm supported at heart height, two readings a minute apart, repeated morning and evening for a week and averaged. After a normal breath out, measure your waist and divide by your height. Walk the 13-foot (4-meter) course twice at your usual pace and take the better time. Beside a counter, do the 30-second chair stand and the ten-second one-legged stance on each leg. Fifteen minutes, and the only purchase in it is a blood pressure monitor.
Then wait. The physical tests are worth repeating about every three months, waist about once a season, and the week of blood pressure about once a year unless a prescriber wants more. If tracking any of these makes you check more often rather than less, that is the signal to stop. The measurement is meant to inform one decision and then stop.
Go Deeper
- Grip strength: the dynamometer protocol, age and sex reference ranges, and how to raise the number.
- Walking speed: the sixth vital sign, what each speed band means, and how to move it.
- Balance and fall prevention: the chair-stand and balance cut-offs, and the training that actually lowers fall risk.
- Sleep regularity: why steady timing beats chasing hours, and how to steady it.
- VO2max training: the fitness measure you cannot take at home, and how to raise it.
- High blood pressure: what a home reading means, and what actually moves it.
The Chinese Medicine View
The tradition has always read the body's own signs, the tongue, the pulse, the complexion, as a window on the whole person, and it measures something different from any number on this page. Nothing in it produces a figure, and nothing in it forecasts a decade. What a person is asked to notice belongs mostly to the asking pillar of the four examinations: sleep and how it breaks, appetite and digestion, thirst, stool and urine, sweating and when it happens, where energy sits through the day, whether cold or heat troubles you more, and in women the cycle. Our four pillars: asking page carries the full set, and looking covers what the practitioner observes.
Tongue observation you can do yourself, in a mirror, first thing, before coffee, tea, toothpaste or scraping, in daylight, with the tongue out and relaxed. It is good for noticing change in yourself over weeks. It is not a self-diagnosis, and the reason is measured: the formal inter-rater agreement between practitioners reading the same tongue is limited, moderate for coating and fair for body color in the studies above. That is a finding about standardization, and both studies place the cause in the definitions rather than the seeing, since agreement rose once practitioners used an explicit classification scheme.
The limited reliability score does not erase the clinical value a skilled practitioner draws from the whole picture, and a reliability number is not a verdict on the whole art. It is a reason to read your own tongue as a trend rather than an answer, and it is why tongue reading sits off your home checklist. Our tongue diagnosis page teaches the full system.
Pulse diagnosis is not a self-measure. Taking your own radial pulse in three positions at three depths while relaxed enough for the reading to mean anything is not something a person does to themselves, and the reliability literature is limited even among practitioners, with acceptable agreement only where the method carried concrete operational definitions. Counting your own rate is useful and is a different act from pulse palpation.
The tradition also holds a person back from over-reading a single sign. A pattern is read from the convergence of several signs, so one sign read alone is a fragment, and the classical texts treat it that way. The Neijing's account of nourishing life puts the emphasis on regulating sleep, food, work, emotion and the seasons rather than on inspecting yourself, and constant self-examination is its own disturbance to the Shen. For someone anxious, that caution applies directly: adding a daily tongue check to a day already spent scanning the body makes the day worse. Strength of tradition does not by itself establish a mechanism, and a reliability score is not a verdict on what a skilled hand can read. Both lenses hold at once.
Common Questions
What is the single best health test I can do at home?
Walking speed over 13 feet (4 meters) at your usual pace. Every 0.1 m/s faster came with about 12% lower mortality across 34,485 older adults, predicting survival about as well as a model built from age, sex, chronic conditions, smoking, blood pressure, body mass index and hospitalization together. It takes seconds and needs a stopwatch and a floor.
If I practice these tests, will my health improve?
Your score will improve. Whether the risk improves depends on the test. Progressive resistance training raised gait speed by 0.08 m/s and improved chair-stand performance moderately to largely, and improvement in walking speed over a year did predict better eight-year survival. For grip strength and the balance tests, nobody has shown that raising the score moves the outcome it was reading.
Why is my blood pressure different at the doctor's office?
Partly the setting and partly the technique. An unsupported arm reads 6.5 mmHg high, a hand in your lap 3.9 mmHg high, and a cuff too small for the arm up to 19.5 mmHg high. When home and office readings were entered into the same model in 2,081 Finnish adults, only the home reading still predicted cardiovascular events.
Are smartwatch sleep scores accurate?
For how long you slept, reasonably. For sleep stages, no. Against polysomnography in 34 adults, consumer devices detected sleep with sensitivity of 0.93 or better but caught wake with specificity of 0.18 to 0.54, staging was inconsistent between devices, and all of them did worse on disrupted nights. Day-to-day changes in total sleep time and bedtime are the parts worth reading.
Should I get a full panel of blood tests to check everything?
Testing more found more and changed nothing. Across 17 randomized trials and 251,891 participants, general health checks made no difference to total mortality on high-certainty evidence. A specific test asked for a specific reason is a different thing entirely, and that conversation belongs with a doctor who knows your history.
Explore Related
Other pages this one connects to, by the evidence they share, the outcomes they touch, and the ground they cover.
How this connects
- Related
-
- Resistance Training
- Diagnostic Lab Panels · The free tier under the paid one.
Pages that lead here: Balance & Fall Prevention
All 32 sources on this page independently checked and cross-referenced.
Thomas Dehli, Founder & Editor, Sacred Lotus
Sacred Lotus has published Chinese medicine reference material since 2001. Integrative pages are held to the same standard as the herb and formula library: cite the source, grade the claim at its real strength, and say where the research has not looked. This page is educational and it is not medical advice. Last reviewed and updated August 10, 2026.
Evidence strength
How confidently the research supports a claim. Strength describes the evidence, not our endorsement.