PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

Cross-Validation and Normalization in EEG Sleep Staging: Impacts on Generalization, Calibration, and Clinical Validity

AKAhmet Sertol Köksal

Key Points

  • Sleep staging accuracy significantly improves with subject-aware normalization and test-time adaptation.
  • The record-wise Macro-F1 score was higher at 0.70, but subject-wise evaluations showed a decrease of 9 percentage points.
  • Normative practices differ; adopting subject-wise and leave-one-subject-out evaluation provides better generalization.
  • Calibrated predictions through effective normalization support safer clinical decisions and reduce misclassifications.

Abstract

Accurate and reliable sleep staging from electroencephalography (EEG) is essential for both research and clinical applications. However, evaluation practices differ widely, and subtle methodological choices can strongly influence reported results. In this study, we examined how cross-validation strategies and normalization protocols affect the reliability and generalizability of EEG-based sleep staging models. Two benchmark datasets, SleepEDF and ISRUC, were used to systematically compare common approaches. We found that record-wise evaluation, often used in the literature, leads to overly optimistic results, while subject-wise and leave-one-subject-out (LOSO) evaluations provide more realistic estimates. On SleepEDF and ISRUC, record-wise median Macro-F1 was 0.70 and 0.71, respectively; under subject-wise it was lower by 9 and 7 percentage points. Similarly, normalization strategies matter: although fold-aware normalization performed better in standard tests, subject-aware normalization combined with test-time adaptation produced the most consistent and clinically relevant outcomes, which improves calibration (lower ECE) and supports safer decisions. In particular, it reduced errors and improved both classification accuracy and probability reliability; for example, on ISRUC, subject-aware further improved Macro-F1 by 0.08, reduced ECE by 0.02, and increased kappa by 0.10, compared with fold-aware normalization. We present a protocol-level, model-independent proof that evaluation and normalization decisions can compete with model selection, particularly when datasets change. Better-calibrated predictions and safer clinical decisions are obtained by using subject-wise/LOSO for internal assessment and subject-aware normalization with test-time adaptation for deployment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ahmet Sertol Köksal (2026) studied this question.

synapsesocial.com/papers/69a75c89c6e9836116a257a3https://doi.org/10.29130/dubited.1773372
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Estimating distribution shifts for predicting cross-subject generalization in electroencephalography-based mental workload assessment2022 · 9 citations
  2. 2Intra- and Inter-subject Variability in EEG-Based Sensorimotor Brain Computer Interface: A Review2020 · 319 citations
  3. 3Coupled electrophysiological, hemodynamic, and cerebrospinal fluid oscillations in human sleep2019 · 1,226 citations
  4. 4Alternative Electrode Placement in (Automatic) Sleep Scoring (F pz-Cz/P z-Oz versus C4-At )1990 · 72 citations
  5. 5ZleepAnlystNet: a novel deep learning model for automatic sleep stage scoring based on single-channel raw EEG data using separating training2024 · 17 citations