Content extraction from HTML documents

  • Rahman A
  • Alam H
  • Hartono R
N/ACitations
Citations of this article
35Readers
Mendeley users who have this article in their library.

Abstract

In recent times, the way people access information from the web has undergone a transformation. The demand for information to be accessible from anywhere, anytime, has resulted in the introduction of Personal Digital Assistants (PDAs) and cellular phones that are able to browse the web and can be used to find information using wireless connections. However, the small display form factor of these portable devices greatly diminishes the rate at which these sites can be browsed. This shows the requirement of efficient algorithms to extract the content of web pages and build a faithful reproduction of the original pages with the important content intact.

Cite

CITATION STYLE

APA

Rahman, A. F. R., Alam, H., & Hartono, R. (2001). Content extraction from HTML documents. Proceedings of the First International Workshop on Web Document Analysis ( WDA 2001), (July), 7–10.

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free