Abstract
Automatic sleep staging with deep learning has advanced considerably, yet clinical adoption remains hindered by limited generalization, model bias, and inconsistent evaluation practices. We present SLEEPYLAND, an open-source framework comprising ~ 220,000 h of in-domain and ~ 84,000 h of out-of-domain polysomnographic recordings, spanning diverse ages, disorders, and hardware configurations. We release pre-trained state-of-the-art models, evaluating them across single- and multi-channel EEG/EOG setups. We introduce SOMNUS, an ensemble that integrates models via soft-voting, achieving robust performance across 24 datasets (macro-F1, 68.7–87.2%), outperforming individual models in 94.9% of cases and exceeding prior state-of-the-art. Exploiting the Bern-Sleep-Wake-Registry (N = 6633), we show that while SOMNUS improves generalization, no model architecture consistently minimizes model demographic/clinical bias. On multi-annotated datasets, SOMNUS surpasses the best human scorer (macro-F1, 85.2% vs 80.8% on DOD-H, and 80.2% vs 75.9% on DOD-O), more closely reproducing consensus. Finally, ensemble disagreement metrics predict scorer ambiguity (ROC-AUC 82.8%), providing reliable proxies for human uncertainty.
Cite
CITATION STYLE
Rossi, A. D., Metaldi, M., Bechny, M., Filchenko, I., Meer, J. van der, Schmidt, M. H., … Fiorillo, L. (2026). SLEEPYLAND: trust begins with fair evaluation of automatic sleep staging models. Npj Digital Medicine, 9(1). https://doi.org/10.1038/s41746-025-02237-2
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.