High-accuracy splice site prediction based on sequence component and position features

26Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.

Abstract

Identification of splice sites plays a key role in the annotation of genes. Consequently, improvement of computational prediction of splice sites would be very useful. We examined the effect of the window size and the number and position of the consensus bases with a chi-square test, and then extracted the sequence multi-scale component features and the position and adjacent position relationship features of consensus sites. Then, we constructed a novel classification model using a support vector machine with the previously selected features and applied it to the Homo sapiens splice site dataset. This method greatly improved cross-validation accuracies for training sets with true and spurious splice sites of both equal and different proportions. This method was also applied to the NN269 dataset for further evaluation and independent testing. The results were superior to those obtained with previous methods, and demonstrate the stability and superiority of this method for prediction of splice sites. © FUNPEC-RP www.funpecrp.com.br.

Cite

CITATION STYLE

APA

Li, J. L., Wang, L. F., Wang, H. Y., Bai, L. Y., & Yuan, Z. M. (2012). High-accuracy splice site prediction based on sequence component and position features. Genetics and Molecular Research, 11(3), 3431–3451. https://doi.org/10.4238/2012.September.25.12

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free