Método de mineração de dados para identificação de câncer de mama baseado na seleção de variáveis

5Citations
Citations of this article
24Readers
Mendeley users who have this article in their library.

Abstract

In the majority of countries, breast cancer among women is highly prevalent. If diagnosed in the early stages, there is a high probability of a cure. Several statistical-based approaches have been developed to assist in early breast cancer detection. This paper presents a method for selection of variables for the classification of cases into two classes, benign or malignant, based on cyto-pathological analysis of breast cell samples of patients. The variables are ranked according to a new index of importance of variables that combines the weighting importance of Principal Component Analysis and the explained variance based on each retained component. Observations from the test sample are categorized into two classes using the k-Nearest Neighbor algorithm and Discriminant Analysis, followed by elimination of the variable with the index of lowest importance. The subset with the highest accuracy is used to classify observations in the test sample. When applied to the Wisconsin Breast Cancer Database, the proposed method led to average of 97.77% in classification accuracy while retaining an average of 5.8 variables.

Cite

CITATION STYLE

APA

Holsbach, N., Fogliatto, F. S., & Anzanello, M. J. (2014). Método de mineração de dados para identificação de câncer de mama baseado na seleção de variáveis. Ciencia e Saude Coletiva, 19(4), 1295–1304. https://doi.org/10.1590/1413-81232014194.01722013

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free