Comparative exploration of document collections: A visual analytics approach

30Citations
Citations of this article
70Readers
Mendeley users who have this article in their library.

Abstract

We present an analysis and visualization method for computing what distinguishes a given document collection from others. We determine topics that discriminate a subset of collections from the remaining ones by applying probabilistic topic modeling and subsequently approximating the two relevant criteria distinctiveness and characteristicness algorithmically through a set of heuristics. Furthermore, we suggest a novel visualization method called DiTop-View, in which topics are represented by glyphs (topic coins) that are arranged on a 2D plane. Topic coins are designed to encode all information necessary for performing comparative analyses such as the class membership of a topic, its most probable terms and the discriminative relations. We evaluate our topic analysis using statistical measures and a small user experiment and present an expert case study with researchers from political sciences analyzing two real-world datasets. © 2014 The Eurographics Association and John Wiley & Sons Ltd. Published by John Wiley & Sons Ltd.

Cite

CITATION STYLE

APA

Oelke, D., Strobelt, H., Rohrdantz, C., Gurevych, I., & Deussen, O. (2014). Comparative exploration of document collections: A visual analytics approach. Computer Graphics Forum, 33(3), 201–210. https://doi.org/10.1111/cgf.12376

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free