Abstract
The paper aims at designing a scheme for automatic identification of a species from its genome sequence. A set of 64 three-tuple keywords is first generated using the four types of bases: A, T, C and G. These keywords are searched on N randomly sampled genome sequences, each of a given length (10,000 elements) and the frequency count for each of the 4(3) = 64 keywords is performed to obtain a DNA-descriptor for each sample. Principal Component analysis is then employed on the DNA-descriptors for N sampled instances. The principal component analysis yields a unique feature descriptor for identifying the species from its genome sequence. The variance of the descriptors for a given genome sequence being negligible, the proposed scheme finds extensive applications in automatic species identification. An alternative approach to automatic species classification and identification of species using Self-Organizing Feature Map is also discussed in the paper. The computational map is trained by using the DNA-descriptors from different species as the training inputs. The maps for different dimensions are constructed and analyzed for optimum performance. The scheme presents a novel method for identifying a species from its genome sequence with the help of a two dimensional map of neuronal clusters, where each cluster represents a particular species. The map is shown to provide an easier technique for recognition and classification of a species based on its genomic data.
Author supplied keywords
Cite
CITATION STYLE
Sen, S., Narasimhan, S., & Konar, A. (2007). Biological data mining for genomic clustering using unsupervised neural learning. Engineering Letters, 14(2). Retrieved from http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.148.4843&rep=rep1&type=pdf
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.