Abstract
The evolution of the World Wide Web has created a great opportunity for data production and for the construction of public repositories that can be accessed all over the world. However, as our ability to generate new data grows, there is a dramatic increase in the need for its efficient integration and access to all the dispersed data. In specific fields such as biology and biomedicine, data integration challenges are even more complex. The amount of raw data, the possible data associations, the diversity of concepts and data formats, and the demand for information quality assurance are just a few issues that hinder the development of a general proposal and solid solutions. In this article we describe a lightweight information integration architecture that is capable of unifying, in a single access point, several heterogeneous bioinformatics data sources. The model is based on web crawling that automatically collects keywords related with biological concepts that are previously defined in a navigation protocol. This crawling phase allows the construction of a link-based integration mechanism that conducts users to the right source of information, keeping the original interfaces of available information and maintaining the credits of original data providers.
Author supplied keywords
Cite
CITATION STYLE
Lopes, P., Arrais, J., & Oliveira, J. L. (2009). Link integrator: A link-based data integration architecture. In KDIR 2009 - 1st International Conference on Knowledge Discovery and Information Retrieval, Proceedings (pp. 274–277). https://doi.org/10.5220/0002291702740277
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.