Abstract
Availability of high through put gene expression data has enabled computational analysis of it for early diagnosis of diseases like cancer. This data contains expression values of thousands of genes in the genome of an organism. However, this gene expression data is very high dimensional, one dimension each corresponding to one genes in the genome and very few of these genes are associated with a disease. At the same time, the number of samples or observations available is very small as compared to the number of features, also this data suffers from class imbalance. Therefore, the task of selecting the genes that are relevant to the disease being studies is an important task and being researched widely in the computational sciences. In this paper, we have proposed a randomized ensemble method for feature selection from cancer gene expression data using a combination of mutual information and recursive feature elimination. The approach has been applied on Leukemia gene expression dataset. We obtained a classification accuracy of 99% with a gene subset of size 316 genes and with a subset of size 4 the accuracy is 95%. Thus we achieved a dimensionality reduction of 98.5% with 99% accuracy. Comparison with standard methods shows that the proposed method performs better.
Author supplied keywords
Cite
CITATION STYLE
Koul, N., & Manvi, S. S. (2020). Ensemble Feature Selection from Cancer Gene Expression Data using Mutual Information and Recursive Feature Elimination. In Proceedings of 2020 3rd International Conference on Advances in Electronics, Computers and Communications, ICAECC 2020. Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/ICAECC50550.2020.9339518
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.