BACKGROUND: When LDL cholesterol (LDL-C) estimation equations are benchmarked against "direct" LDL-C (LDLdirect) values from electronic health records (EHRs), the implicit assumption is that reported direct values derive from chemical assays rather than from calculations. This assumption has not been systematically tested. METHODS: We analyzed 19 470 same-day lipid panels from the All of Us Research Program (50+ health provider organizations). We defined the effective divisor, Deff = TG/(TC - HDL-C - LDLdirect), where TG is triglycerides and TC is total cholesterol, and tested for evidence that reported "direct" values were Friedewald-calculated (Deff ≈ 5.0), with validation on the National Health and Nutrition Examination Survey (NHANES) and the Medical Information Mart for Intensive Care IV (MIMIC-IV). RESULTS: Of 19 470 panels (from All of Us), 7.6% (1482) were algebraically confirmed Friedewald-substituted (|LDLFW - LDLdirect| ≤ 0.02 mg/dL 0.0005 mmol/L, where LDLFW is the Friedewald-calculated LDL-C) and 31.1% met a broader criterion (≤ 0.5 mg/dL 0.013 mmol/L). The Deff distribution showed a 27.6-fold spike at 5.0 (z = 68 vs empirical background; z = 263 vs shuffled-LDL null). Contamination deflated observed Friedewald mean absolute error (MAE) by 30.7% (11.85 vs 17.10 mg/dL 0.306 vs 0.442 mmol/L; true MAE 44.3% higher). Near treatment thresholds, contamination masked 12 percentage points of misclassification among near-threshold cases at 70 mg/dL 1.81 mmol/L (26% apparent vs 38% true). Within-patient discordance confirmed substitution (39.6-fold gap, P < 10⁻⁵). NHANES validation showed 28.3-fold enrichment. In MIMIC-IV, the spike was absent in measured panels (0.8×) but present in calculated panels (z = 210). CONCLUSIONS: A minimum of 7.6% of "direct" LDL-C values in this national multi-site EHR cohort are algebraically confirmed Friedewald-substituted (up to 31.1% under a broader criterion), deflating observed equation error and attenuating the apparent advantage of improved methods. Deff decontamination should be a standard preprocessing step for EHR-based equation benchmarking.
Doku et al. (Tue,) studied this question.