Skip to main content

← Longevity Study · Explore

The Data Paradox

Income marks a 2.3× mortality gap. The available sleep-quality items also vary with health outcomes. Yet adding either feature set changes these predictive models very little after measured health is included.

The Evidence Is Real

Income and sleep quality are both associated with health outcomes. The question here is narrower: does adding these measurements improve prediction beyond the variables already in each model?

Income & Mortality

2.3× mortality gap

Poorest income quintile: 28.0% die within 10 years. Highest: 12.3%. Data from 28,636 adults pooled across ten NHANES cycles (1999–2018).

If wealth predicts death this strongly, shouldn't it improve our models?

Sleep & Health Decline

69% vs 39% outcome rate

The corrected RAND HRS coding compares five self-reported sleep-quality indicators among 24,155 people. The outcome rates are rendered directly from the promoted aggregate data below.

If sleep quality separates outcomes this clearly, shouldn't it help?

The ROC Curves Tell the Truth

ROC curves reveal how well a model distinguishes positive from negative cases across every possible threshold. If adding data helps, the curve should shift upward. Watch what actually happens.

Adding Income & Education Data

Q15: 28,636 NHANES adults · 6,626 deaths · 10-year follow-up

Adding Sleep Data

Q16: 24,155 people · 6-year health decline outcome

No gain for the domain score: Adding the available sleep-quality variables changed the domain model from AUC 0.604 to 0.599. This held-out comparison does not explain why performance changed.

The Reclassification Shell Game

Net Reclassification Index (NRI) counts how many patients move to a more appropriate risk category. For every person the new data helps, another is hurt.

SDOH Reclassification

28,636 adults across 4 risk categories

Sleep Reclassification

24,155 patients across 4 risk categories

The Signal Is Already Captured

Feature importance reveals why: the new variables rank near the bottom. Age, self-rated health, and existing conditions already carry the signal that wealth and sleep correlate with.

Health + SDOH Model Features

GradientBoosting importance — SDOH features highlighted

Health + Sleep Model Features

GradientBoosting importance — sleep feature highlighted

Why SDOH Adds Little Here

Income, education, and marital status overlap predictively with measured health in this cohort. The analysis does not determine whether that overlap reflects causal pathways, reverse causation, shared causes, selection, or the chosen outcome window.

Why Sleep Doesn't Help

In this cohort, the available sleep-quality items overlap predictively with health, depression, and function measures already in the model. The comparison cannot determine whether that overlap reflects causation, reverse causation, shared causes, or measurement limits.

Two Paradoxes, One Lesson

The longevity study surfaced two symmetrical findings. Together, they define the boundaries of the right-fidelity principle.

The Biology Paradox

More computation doesn't always help. ML alone (r = 0.478) was worse than a textbook formula (r = 0.527) for biological age.

The Data Paradox

More data doesn't always help. Adding SDOH (+0.003 AUC) and sleep (+0.002 AUC) to these models changed discrimination very little.

The ADM Principle

The right model at the right fidelity. Not the most complex, not the most data-rich — the one matched to the question and the decision it supports.

The right-fidelity lesson isn't just "simpler is better." In 7 of 13 longevity investigations, ML genuinely outperformed domain knowledge — sometimes dramatically (diabetes risk on NHANES: AUC 0.7840.851, +6.7 points). The lesson is that fidelity has a ceiling for each question, and adding complexity beyond that ceiling wastes resources without improving decisions. Analysis Driven Modeling™ finds that ceiling before you build.