Abstract
Tandem mass spectrometry (MS/MS) combined with protein database searching has been widely used in protein identification. A validation procedure is generally required to reduce the number of false positives. Advanced tools using statistical and machine learning approaches may provide faster and more accurate validation than manual inspection and empirical filtering criteria. In this study, we use two feature selection algorithms based on random forest and support vector machine to identify peptide properties that can be used to improve validation models. We demonstrate that an improved model based on an optimized set of features reduces the number of false positives by 58% relative to the model which used only search engine scores, at the same sensitivity score of 0.8. In addition, we develop classification models based on the physicochemical properties and protein sequence environment of these peptides without using search engine scores. The performance of the best model based on the support vector machine algorithm is at 0.8 AUC, 0.78 accuracy, and 0.7 specificity, suggesting a reasonably accurate classification. The identified properties important to fragmentation and ionization can be either used in independent validation tools or incorporated into peptide sequencing and database search algorithms to improve existing software programs. © 2008 Imperial College Press.
Author supplied keywords
Cite
CITATION STYLE
Fang, J., Dong, Y., Williams, T. D., & Lushington, G. H. (2008). Feature selection in validating mass spectrometry database search results. In Journal of Bioinformatics and Computational Biology (Vol. 6, pp. 223–240). https://doi.org/10.1142/S0219720008003345
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.