Forecasting Later Cognitive Classification.
Calibrated 2-year forecasting on the Health and Retirement Study, measured against the Langa-Weir silver-label proxy — a research classification, not a clinical diagnosis. The complex models’ point estimates did not exceed the operating threshold, but their uncertainty intervals crossed it, so the decision to stop at the simple rung is provisional and specific to this benchmark.
By Michael Key · ORCID
The cohort, harmonized variables, and cognitive outcome belong to others. The Health and Retirement Study (HRS) is sponsored by the National Institute on Aging (NIA grant U01AG009740) and conducted by the University of Michigan. The longitudinal variables come from the RAND HRS Longitudinal file. The outcome — Normal / Cognitive Impairment No Dementia (CIND) / Dementia — is the Langa-Weir classification, a research proxy rather than a clinical diagnosis. These are registered public releases governed by the HRS Conditions of Use; raw records are not redistributed here. This study’s contribution is not new data or a new measurement. It is a calibrated evaluation against that silver label. This is not a clinical diagnostic, individual advice, or a statement about any person.
What this means. In this HRS benchmark, the minimal age-and-cognition model produced reasonably calibrated probabilities to the Langa-Weir research label. Adding features and model complexity produced small gains, while accounting for death and checking subgroups materially changed how the results should be read.
What it does not mean. It is not a clinical dementia forecast, it is not externally validated, and it does not prove that the simple model is equivalent to the more complex models or sufficient for another use.
Three limits govern every result.
1. Silver label. Every metric is measured against the Langa-Weir classification — an algorithmic research proxy for dementia. One ADAMS comparison reported 52% overall sensitivity for the Langa-Weir dementia classification, with a large self/proxy split (24% among self-respondents and 82% among proxy respondents); a later validation found sensitivities of 18–62% across five HRS algorithms, with subgroup variation (Crimmins et al. 2011; Gianattasio et al. 2019). Published evaluations also document race/ethnicity differences in sensitivity, specificity, and accuracy that can bias disparity estimates. With no ADAMS ground truth in this study, measured disparities cannot be decomposed into label bias, measurement differences, sampling, case mix, and biology. Read every “calibrated” and every AUROC here as to the silver label.
2. Self-respondent selection. The primary analysis keeps only people still self-respondents at the next interview — and decline drives the switch to a proxy. The self-respondent share at baseline falls from 0.962 (Normal) to 0.575 (Demented); among self-respondents re-interviewed, the self→proxy transition rate by next wave climbs from 0.013 (Normal) to 0.178 (Demented). The self→proxy filter alone dropped 2,233 person-periods carrying 1,282 imputed-dementia events. These estimates do not transport directly to the full at-risk population. Observed incidence is likely biased downward; discrimination and calibration could shift in either direction. An imputed-inclusive run is a second directional sensitivity, not a bound on the truth: its imputation model uses adjacent-wave cognition scores and may carry entry-wave future information.
3. Pooled AUROC is separation-inflated. The pooled AUROC is lifted by the model separating Normal from CIND — a baseline-status feature, not foresight into who will convert. Within-stratum discrimination is the more relevant test of conversion within each baseline state; every pooled AUROC is labeled and paired with those strata where available.
A minimal cognition model whose tested escalations did not clearly earn the added fidelity.
The model and the holdout
A discrete-time pooled-logistic on age, cognitive score, and lagged baseline status, evaluated on an out-of-time holdout (W15 2020 → W16 2022): 5,022 person-periods, 195 incident events, base rate 0.039.
Calibration first
The minimal cognition model is reasonably calibrated to the silver label — calibration slope 0.89 on the temporal holdout. ECE is 0.0097, though at a base rate of 0.039 ECE is mechanically small; the slope is the lead metric. Person-grouped cross-validation gives slope 1.00, ECE 0.0011, and pooled AUROC 0.873. That supports stability across these folds; it is not an external validation and does not remove dependence among a person’s repeated observations.
Discrimination: pooled is optimistic, within-stratum is honest
Pooled AUROC on the temporal holdout: 0.850 [0.821, 0.880]. This sits at the optimistic top of the analog band in a recent external-validation review — pooled c-statistics in the low-to-mid 0.70s for the best-validated scores, reaching into the low 0.80s for the strongest EHR models (Stephan et al. 2026, BMC Medicine) — and it is separation-inflated (caveat 3). Within-stratum discrimination is 0.733 (Normal) / 0.675 (CIND). The lower value in CIND shows that discrimination is weakest in the higher-risk stratum nearest the classification threshold.
Why the simple model is the right fidelity for this decision
Two model-complexity escalations were tested against the minimal cognition rung, each with a paired Δ-AUROC on the same holdout rows and a TOST equivalence test. The Δ = 0.02 operating threshold was adopted for the review-driven paired reanalysis; it was not a conventional preregistration. A separate recalibration rung was evaluated in probability space using ECE and Brier score:
- Penalization (elastic-net). The fuller feature set adds a paired Δ-AUROC of +0.0196 [0.011, 0.028]. The point estimate sits just below 0.02, but the interval includes gains above the threshold and TOST is inconclusive at 195 events.
- Gradient boosting (HistGBM). Paired Δ-AUROC +0.018 [0.005, 0.032]. Its point estimate also sits below the threshold, but the interval spans it. GBM and the tuned linear full models have overlapping intervals; that is not proof of equivalence.
- Recalibration. The raw model is already well-calibrated (slope 0.89). Platt recalibration is a wash, and isotonic adds no decision-relevant benefit by ECE or Brier score. A pooled recalibrator also cannot fix the mild within-Normal over-confidence.
The ADM conclusion is provisional and decision-specific. The complex models’ point estimates are below the 0.02 operating threshold, and recalibration adds no decision-relevant benefit by the reported ECE or Brier measures. For this benchmark, that is enough to stop at the minimal model. But the paired intervals span the threshold and TOST is inconclusive: this study does not establish equivalence or prove that the extra fidelity is too small to matter elsewhere.
- Self-respondent selection (caveat 2 applies here). The self→proxy filter drops 2,233 person-periods carrying 1,282 imputed-dementia events; the highest-risk converters exit the observed primary earliest. The estimates do not transport directly to the full population, and the direction of discrimination or calibration bias is not established.
- Riley power caveat. The primary 195-event holdout sits just under the Riley ≥200/≥200 calibration-sample rule (
riley_ok=False). The ≥50 and imputed-inclusive variants (248 / 329 events) corroborate every direction. - Pooled AUROC is separation-inflated (caveat 3). The 0.850 pooled figure reflects Normal-vs-CIND separation. The 0.733 / 0.675 within-stratum numbers are the ones to trust.
- All metrics are to the silver label (caveat 1). No ADAMS ground truth is available to decompose label-bias from biology.
Death is a competing risk, and it moves the absolute number.
A naïve estimator that censors death as if it were random overstates cumulative incidence — this is a known identity from competing-risks theory, quantified here for this cohort. The overstatement scales with the competing death rate, and death competes hardest for exactly the people most at risk of dementia.
Two-year mortality in the CIND stratum is 0.136 versus 0.052 in Normal — a more than two-fold difference. The practical consequence:
- Overall: naïve incidence 0.0319 → competing-risk 0.0297, a relative overstatement of ~7%.
- CIND stratum: 0.1343 → 0.1160, a relative overstatement of ~14% — because death competes hardest where dementia risk is highest.
The discrete-time multinomial that models death as an explicit third outcome has pooled AUROC 0.856 (separation-inflated), within-stratum AUROC 0.768 Normal / 0.693 CIND, and internal calibration to the silver label. These are cohort-internal estimates, not externally validated clinical risks.
The load-bearing takeaway is the relative difference between estimators in this cohort, not the absolute level. The absolute level carries the self-respondent-selection bias of caveat 2. Censoring death moves the estimated silver-label incidence upward, most in the CIND stratum.
This is a known competing-risks identity, quantified for this cohort — not a new discovery. The “shift ≈ death-rate identity” flag in the analysis is a sanity check, not independent evidence. Self-respondent selection likely biases the absolute incidence downward (caveat 2), and neither the absolute nor relative estimate has been externally validated.
Aggregate calibration can conceal subgroup-performance differences under an uneven silver label.
The minimal cognition model includes age but excludes race, ethnicity, sex, and education features. In aggregate it has out-of-fold slope 0.996, ECE 0.0010, and pooled AUROC 0.873 [0.867, 0.880] (separation-inflated). Excluding those fields and looking good in aggregate do not establish fairness. Stratified, the aggregate conceals disparity:
| Subgroup | Silver-label base rate | AUROC (pooled) | Cal. slope |
|---|---|---|---|
| White, non-Hispanic | 0.0228 | 0.882 [0.874, 0.890] | 1.021 |
| Black, non-Hispanic | 0.0692 | 0.820 [0.805, 0.835] | 0.857 |
| Hispanic | 0.0622 | 0.830 [0.808, 0.848] | 0.917 |
| College or more † | 0.0096 | 0.903 [0.880, 0.922] | 1.072 |
| Less than high school | 0.0852 | 0.804 [0.793, 0.816] | 0.862 |
† Riley-underpowered (179 events < 200); read this interval as wide. All AUROCs are pooled-across-strata (caveat 3). Lead with slope and AUROC, not raw ECE — ECE’s cross-group ratio is base-rate-inflated.
The calibration-slope point estimates indicate greater over-confidence for Black, Hispanic, and less-educated respondents (0.857–0.917 for Black/Hispanic/0.862 <HS vs 1.021 White / 1.072 College+), and the model is less discriminating for those same groups, with non-overlapping White-vs-Black and College-vs-<HS AUROC intervals. The slope estimates were not bootstrapped, so their cross-group differences are descriptive rather than formal tests. The model excludes race, ethnicity, sex, and education features, but age, cognitive measures, sampling, case mix, and the target label can still carry structural differences.
Does case mix explain the gap? Not fully in this sample. The pooled subgroup AUROCs are separation-inflated and every group falls sharply within-stratum, so case mix explains part of the pooled gap. Differences also remain within the cognitively normal stratum. Those cells are thin and most fail the Riley threshold; they are indicative, not definitive, and support no causal attribution.
Several possible failure modes remain entangled. Without ADAMS ground truth here, the slope gaps cannot be separated into label bias, measurement differences, sampling, case mix, and biology. The race and education base-rate gaps (~3x by race, ~9x by education) are not established true-incidence differences.
The continuous score is forecastable — with a pattern consistent with partial de-noising.
Beyond binary dementia crossing: can the next-wave 27-point cognitive score be forecast? A Ridge regression on cognition-only features lowers MAE from 3.031 (persistence baseline) to 2.690 (cognition-only); the full feature set reaches 2.611; GBM-regressor 2.616 ≈ Ridge. All model confidence intervals sit below persistence.
The honest deflation. The gain is concentrated in the Normal stratum, where MAE falls from 3.070 under persistence to 2.663 with cognition-only. Cognition-only ties persistence in the at-risk CIND group at age ≥65 (2.829 vs 2.831) — only the full feature set extracts a small CIND gain (2.687). That pattern is consistent with some de-noising or mean reversion in a noisy, autocorrelated test score, but the study did not decompose the error into those components. GBM ≈ Ridge; the full feature set adds little beyond cognition-only in this benchmark.
Measured only on respondents who stayed self-respondents at t+1, selecting against the steepest decliners (caveat 2). This limits transport to the full at-risk population; the direction of the error-metric bias is not established.
What this study cannot establish.
Out-of-cohort generalisation is not claimed. Harmonised ELSA/CHARLS cross-cohort files are not on disk; out-of-cohort transfer would need them acquired and Langa-Weir-mapped — buildable, not done, stated as a limitation rather than softened. Separately and more fundamentally, no ADAMS clinical ground truth is on disk — label bias and biology cannot be decomposed. Two different limits: a transfer-generalisation gap and a label-validity gap.
The t+2 (~4-year) horizon is deferred, not reported. It needs its own 4-year competing-risk and informative-attrition treatment; the substrate is horizon-parameterized so it remains a flagged future excursion.
None of the tested escalations earned a change in fidelity under the selected operating rule; formal equivalence was not established. The complex models’ point estimates are below 0.02, but their confidence intervals span the threshold. Recalibration adds no decision-relevant benefit by the reported ECE or Brier measures. The decision is to stop provisionally at the minimal model for this benchmark, not to claim that complexity cannot help.
What this contributes, honestly stated.
- A minimal discrete-time logistic using age, cognition, and lagged status, with reasonably calibrated probabilities of later Langa-Weir silver-label classification — temporal-holdout slope 0.89 and person-grouped cross-validation slope 1.00. This is internal validation, not clinical or external validation.
- A provisional fidelity stop. The complex models’ point estimates sit below the operating threshold, while their intervals span it and formal equivalence is inconclusive. For this benchmark, the selected rule stops at the minimal model; it does not establish sufficiency elsewhere.
- Death-as-competing-risk matters for absolute incidence. Censoring death as random overstates 2-year risk by ~7% overall and ~14% in the high-risk CIND stratum. That relative correction is specific to this cohort; self-respondent selection likely biases the absolute incidence downward.
- Subgroup disparity under an uneven label. The model excludes race, ethnicity, sex, and education features, but calibration-slope point estimates indicate greater over-confidence and AUROC is lower for Black, Hispanic, and less-educated respondents. The slope differences were not formally tested; thin cells and the lack of ADAMS ground truth prevent causal interpretation.
- Every claim is to the silver label. Self-respondent selection limits transport to the full at-risk population and likely biases observed incidence downward. Every displayed result traces to script-written output, and the negatives remain in the record.
This is the mission’s first Validation/Benchmark study on registered public-release data. It establishes an evaluation discipline that later disease-specific work can reuse. This HRS study is an on-ramp, not a clinical tool.
Every displayed result has a trace.
Registered public releases. The RAND HRS Longitudinal file 1992–2022 v1 and Langa-Weir Classification of Cognitive Function 1995–2022, used under the HRS Conditions of Use. The study locks local inputs by sha256 checksum and does not redistribute raw records.
Reproduction. Each of the eight investigations regenerates script-written results. The reproduction harness runs each investigation twice and checks byte-identical output. The code, environment records, protocol, and per-investigation provenance are maintained together and available for review on request while a separate public-release review is pending; this page does not claim that every artifact is independently downloadable today.
Method and comparison source. Discrete-time pooled-logistic hazard model; elastic-net penalization; HistGBM; discrete-time multinomial for competing risks; person-cluster bootstrap confidence intervals; grouped person-CV. The analog band comes from Stephan BCM et al., BMC Medicine 2026. The label caveats are grounded in Crimmins et al. 2011 and Gianattasio et al. 2019.