Abstract
Training machine learning models for tasks such as de novo sequencing or spectral clustering requires large collections of confidently identified spectra. Here we describe a dataset of 2.8 million high-confidence peptide-spectrum matches derived from nine different species. The dataset is based on a previously described benchmark but has been re-processed to ensure consistent data quality and enforce separation of training and test peptides.
Cite
CITATION STYLE
Wen, B., & Noble, W. S. (2024). A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models. Scientific Data , 11(1). https://doi.org/10.1038/s41597-024-04068-4
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.