Why biological-age tests give different answers
Biological-age tests can disagree because they measure different features of aging and compare you with different reference groups. A repeat test can also change without showing a lasting health change. In a 2026 analysis of 18 DNA-based aging measures, many were more consistent when the same sample was measured again than when fresh samples were collected. That distinction matters when interpreting your own result.
An aging clock can carry useful information without revealing one “true age” inside your body. Some clocks predict later health outcomes in groups; some detect changes in studies. Their value depends on the question they were built to answer. A few “years younger” on a report cannot be read as extra years of life. [1] [2] [3] [6] [7] [8]
Research disclosure: in the 2026 reliability paper, R. Sehgal reported consulting fees from LongevityTech.fund, TruDiagnostic and Cambrian BioPharma. [5]
What the evidence supports
Shown: clocks capture patterns associated with age and future health. Their precision and stability vary by method. Some clock changes have also been associated with later survival in observational research. [1] [2] [3] [4] [5] [7]
Not shown: an individual score change does not establish that a treatment chosen to lower it prevents disease or improves function. [6] [8]
What would change the assessment: reliable evidence that a specified change exceeds ordinary variation for that test, followed by trials showing that acting on the result improves health. These are two separate requirements. [4] [5] [8]
First ask what the clock measures
An epigenetic clock is a calculation based on DNA methylation—chemical tags attached to DNA. A laboratory measures these tags at selected sites, and an algorithm combines them into a score. The tags are not little timers. Horvath's original clock was designed to estimate calendar age across many tissues. That demonstrated recognizable age-related methylation patterns, rather than a test for a treatable condition. [1]
A 2015 erratum to Horvath’s paper corrected a calibration error in its cancer-tissue analysis.
Later clocks were built for different tasks. The word “age” can hide those differences:
Output | What the model aims to estimate | What the number means |
|---|---|---|
An age estimate | Calendar age from methylation patterns | Resemblance to patterns in the reference data, with prediction error. [1] |
A risk score expressed in years | GrimAge used information related to mortality | A score associated with risk, not a countdown to death. [2] |
A pace estimate | DunedinPACE estimates long-term physiological change | A rate relative to a reference population, not an age in years. [3] |
An age gap or age acceleration | The difference from a score expected at the same calendar age | A relative position; the exact calculation needs to be specified. [1] [2] [3] |
GrimAge combines methylation estimates of smoking history and selected blood proteins in a model developed to predict mortality, then expresses its output on an age-like scale. DunedinPACE began with repeated measurements of 19 indicators of physiological health over roughly two decades in the Dunedin birth cohort. Researchers then built a blood methylation measure to estimate that pace. [2] [3]
A pace of 0.9 therefore does not mean you are 90% of your calendar age. A GrimAge result five years above your age does not establish that you will die five years earlier. They are different kinds of estimates. [2] [3]
Proteomic clocks use measured blood proteins instead of methylation. Findings for an organ-age model built with one protein assay do not validate a DNA test, an imaging score or another commercial protein panel. A shared name does not make the measurements interchangeable. [10]
Three sources of uncertainty
Imagine a report giving a 52-year-old a biological age of 47.3. This is an illustration, not a real result. The decimal shows how the software prints the answer; it does not show that the test can distinguish a few months of aging.
Technical repeatability asks what happens when the same material is measured again. In a 2022 methods study, some established clocks gave substantially different results on repeat measurements. Changing how the methylation data were combined improved agreement. That finding applies to the methods tested; it does not supply one error margin for all commercial services. [4]
Biological repeatability asks what happens with a fresh sample from the same person. It includes both measurement error and real short-term variation in the body. A change need not indicate a lasting shift in disease risk, disability or survival. [5]
Clinical usefulness asks whether the result justifies an action. Even a perfectly repeatable test would need further evidence before choosing a drug or supplement to lower its score could be expected to improve health. [8]
You may see an intraclass correlation coefficient, or ICC, in a report about reliability. It describes agreement between measurements relative to differences between the people studied. It is neither the percentage of your result that is “correct” nor an error margin in years. A study covering a wide age range can produce a different reliability estimate from one studying people of similar ages. Ask for uncertainty that is relevant to your own comparison. [5]
What the 2026 reliability study found
Sehgal and colleagues reanalyzed existing datasets for 18 methylation measures. They compared repeat measurements of the same samples with samples collected over short intervals around stress, meals, pollution exposure and changes in altitude. Many clocks were technically repeatable; fresh-sample stability was lower. The study did not compare complete commercial testing services under one standardized protocol. [5]
The biological datasets were small and mostly involved younger adults. Meal and stress analyses used the same 34-person dataset. There were no matched unexposed groups sampled at the same times. The findings therefore cannot show that a particular meal or stressor caused every fluctuation. Technical replication was conducted within laboratories using older measurement platforms. [5]
The practical lesson is to question an isolated before-and-after claim. The study does not establish a universal fasting rule, rank brands for every customer or show that all clocks are useless. Both the type of repetition and the people studied matter. [5]
Why results can disagree even without an error
The first difference may be the target: calendar age, mortality-related risk and aging pace answer different questions. The reference group also matters. “Younger than expected” means younger relative to the model's reference patterns. A model developed in one population needs suitable validation before it is applied to another. [1] [2] [3]
The sample can differ too. Blood and saliva contain different biological material. Performance should be documented for the sample actually collected, rather than borrowed from a study using another sample type. [1] [4]
Finally, laboratories can change their measurement platform, processing or algorithm. If the calculation changes, ask whether an earlier sample can be calculated with the same version before interpreting a new score as a personal trend. Different inputs, methods and reference groups mean that disagreement need not indicate an error. Averaging unlike scores does not automatically reveal the truth either; the combined measure would need its own rationale and validation. [1] [4] [10]
Changes can still contain useful information
In a 2026 analysis of 699 InCHIANTI participants, researchers used two or three methylation measurements per person and mortality follow-up of up to 24 years. When models included both the starting score and its change, several clocks showed associations with mortality. Results depended on the clock and statistical model. All participants were of European ancestry, which limits how widely the findings can be generalized. [7]
This is a meaningful observational finding, but it is not a treatment experiment. Underlying health can influence both the clock trajectory and survival. The study did not show that lowering a score reproduces the difference in survival. [7]
The CALERIE trial addressed a different question. Its post hoc methylation analysis—a later analysis of the trial data—included 197 healthy adults without obesity with blood measurements before and during a two-year randomized comparison of calorie restriction with unrestricted eating. DunedinPACE changed modestly, while PCPhenoAge and PCGrimAge—versions designed to reduce measurement noise—did not show significant treatment effects. The result concerns a clock response, not a measured reduction in illness. [4] [6]
A May 2023 correction to the CALERIE paper concerned an author affiliation and did not change the results.
Keeping these results together prevents a misleading claim that “biological age improved” by every measure. Predicting risk, responding to an intervention and serving as a reliable substitute for a health outcome are separate achievements. The FDA's framework for such substitute measurements, called surrogate endpoints, requires evidence for the specific use being proposed. [8]
How Healthy Longevity Clinic experts evaluate the evidence
For someone asking whether a prevention plan is working, the key distinction is between a changed measurement and a changed outlook for health. A shift from an illustrative age score of 47.3 to 45 does not answer that question until the exact clock, ordinary variation and reason for testing are understood. Even a change beyond measurement noise still needs evidence linking an action to a benefit. [4] [5] [8]
The useful clinical conversation separates an exploratory score from a finding that already warrants attention. A favorable clock does not cancel a high blood-pressure reading; US guidance calls for appropriate confirmation outside the office before treating newly identified hypertension. An unfavorable clock alone does not diagnose hidden disease. Symptoms, medical history and established risks remain central to interpreting the report. [9]
Stronger evidence would show that a defined clock-guided plan reduces disease or preserves function compared with care that does not use the clock. Until then, the result can inform a discussion about uncertainty and prevention without setting a target of becoming a certain number of “years younger.” [8]
Make a report easier to interpret
Start with the exact test name and version. Is the headline an age, a pace, an age gap or a percentile? Identify the sample and reference population. Then look for ordinary test-to-test variation and an explanation of what size of individual change can be interpreted.
A statistically detectable change is not necessarily one that matters clinically. If the provider cannot quantify meaningful individual change, a precise-looking percentage does not fill the gap. Nor should an unexpected clock lead automatically to broad additional testing simply to explain the number. A result of “47.3 at age 52” can prompt a useful question about the method; by itself, it cannot justify reducing preventive care or buying a treatment to lower the score further. [4] [5] [8] [9]
When to repeat—and what to keep consistent
Repeating a test can help resolve a defined uncertainty. Repeating it until a pleasing number appears makes interpretation harder. Before another test, ask what change exceeds expected variation for the exact method, whether the interval is justified and what you would do after a higher, lower or unchanged result. These studies do not establish a universal retesting schedule. [4] [5]
If you choose to repeat, use the same sample type, laboratory and calculation where possible, and follow the laboratory's validated collection instructions. Record relevant circumstances. Do not assume everyone should fast, and do not stop prescribed treatment to improve test conditions or a clock result.
An unusually high or low first result may be followed by a less extreme one partly because the first contained chance variation. This is called regression toward the mean. Changing several habits and products at once also prevents an uncontrolled before-and-after comparison from identifying what caused the result. A lower score alone cannot resolve either problem. [6] [7]
Three questions for a consultation
Does this clock estimate age, risk or pace, and has it been validated in people like me using this sample and method?
Is my change larger than ordinary variation, and did the laboratory or algorithm change between tests?
What would we do differently because of this result, and has that decision been shown to improve a health outcome?
Common questions
Which clock should I trust?
Trust a specific claim, rather than a name alone. A clock may be useful for a research association without being ready to guide your treatment. Technical stability and relevance to the health outcome you care about both matter. [4] [5] [8]
Does a lower number mean my treatment worked?
Not by itself. Variation, changes in the method and regression toward the mean can complicate the comparison. Even a reliable change in the clock does not establish that the treatment prevents disease or improves function. [6] [7] [8]
Should I buy several tests and average them?
Different clocks can estimate different things. Averaging age, risk and pace does not produce a validated true age. Ask what the combined measure has been shown to do before using it for a decision. [1] [2] [3]
What remains uncertain
Small and mostly younger samples limit the 2026 short-interval reliability findings; absent matched unexposed groups limit causal interpretation of meals and stress. The InCHIANTI associations depended on the clock and model and involved European-ancestry participants. No universal accuracy percentage, best brand, retesting interval or treatment benefit follows from these studies.
References
- DNA methylation age of human tissues and cell types.
- DNA methylation GrimAge strongly predicts lifespan and healthspan.
- DunedinPACE, a DNA methylation biomarker of the pace of aging.
- A computational solution for bolstering reliability of epigenetic clocks: implications for clinical trials and longitudinal tracking.
- Biological Versus Technical Reliability of Epigenetic Clocks and Implications for Disease Prognosis and Intervention Response.
- Effect of long-term caloric restriction on DNA methylation measures of biological aging in healthy adults from the CALERIE trial.
- Longitudinal changes in epigenetic clocks predict survival in the InCHIANTI cohort.
- Surrogate Endpoint Resources for Drug and Biologic Development.
- Hypertension in Adults: Screening.
- Plasma proteomics links brain and immune system aging with healthspan and longevity.
Disclosure
Prepared with AI assistance. The 2026 reliability paper reports a Systems Age patent application. R. Sehgal reported consulting fees from TruDiagnostic, LongevityTech.fund and Cambrian BioPharma; A. Higgins-Chen reported consulting for TruDiagnostic and FOXO Biosciences. Other cited clock-development and InCHIANTI research also reports patent, licensing, consulting or advisory interests. These disclosures do not by themselves determine the validity of the findings.