Feature selection and classification of MAQC-II breast cancer and multiple myeloma microarray gene expression data

45Citations
Citations of this article
78Readers
Mendeley users who have this article in their library.

Abstract

Microarray data has a high dimension of variables but available datasets usually have only a small number of samples,thereby making the study of such datasets interesting and challenging. In the task of analyzing microarray data for thepurpose of, e.g., predicting gene-disease association, feature selection is very important because it provides a way to handlethe high dimensionality by exploiting information redundancy induced by associations among genetic markers. Judiciousfeature selection in microarray data analysis can result in significant reduction of cost while maintaining or improving theclassification or prediction accuracy of learning machines that are employed to sort out the datasets. In this paper, wepropose a gene selection method called Recursive Feature Addition (RFA), which combines supervised learning andstatistical similarity measures. We compare our method with the following gene selection methods: Support Vector Machine Recursive Feature Elimination (SVMRFE) Leave-One-Out Calculation Sequential Forward Selection (LOOCSFS) Gradient based Leave-one-out Gene Selection (GLGS) To evaluate the performance of these gene selection methods,we employ several popular learning classifiers on the MicroArray Quality Control phase II on predictive modeling (MAQC-II) breast cancer dataset and the MAQC-II multiple myeloma dataset. Experimental results show that geneselection is strictly paired with learning classifier. Overall, our approach outperforms other compared methods. Thebiological functional analysis based on the MAQC-II breast cancer dataset convinced us to apply our method forphenotype prediction. Additionally, learning classifiers also play important roles in the classification of microarray dataand our experimental results indicate that the Nearest Mean Scale Classifier (NMSC) is a good choice due to itsprediction reliability and its stability across the three performance measurements: Testing accuracy, MCC values, and AUC errors. © 2009 Liu et al.

Cite

CITATION STYLE

APA

Liu, Q., Sung, A. H., Chen, Z., Liu, J., Huang, X., & Deng, Y. (2009). Feature selection and classification of MAQC-II breast cancer and multiple myeloma microarray gene expression data. PLoS ONE, 4(12), 1–24. https://doi.org/10.1371/journal.pone.0008250

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free