Introducing biases in document clustering

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

In this paper, we present three criteria for introducing biases in document clustering algorithms, when information characterizing the document collections is available. We focus on collections known to be the result of a document categorization or sample-based document filtering process. Our proposals rely on profiles, i.e., document samples known to have been used for obtaining the collection, to extract statistics which determine the biases to introduce. We conduct an experimental evaluation over a number of collections extracted from the widely used corpus RCV1, which allows us to confirm the validity of our proposals and determine a number of situations where biased clusterings, according to different criteria, outperform their unbiased counterparts.

Author supplied keywords

Cite

CITATION STYLE

APA

Ramírez-Cruz, Y. (2014). Introducing biases in document clustering. Computacion y Sistemas. Instituto Politecnico Nacional. https://doi.org/10.13053/CyS-18-1-2014-024

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free