Abstract:Effective analysis of large text collections remains a challenging problem given the growing volume of available text data. Recently, text mining techniques have been rapidly developed for automatically extracting key information from massive text data. Topic modeling, as one of the novel techniques that extracts a thematic structure from documents, is widely used to generate text summarization and foster an overall understanding of the corpus content. Although powerful, this technique may not be directly applicable for general analytics scenarios since the topics and topic–document relationship are often presented probabilistically in models. Moreover, information that plays an important role in knowledge discovery, for example, times and authors, is hardly reflected in topic modeling for comprehensive analysis. In this paper, we address this issue by presenting a visual analytics system, VISTopic, to help users make sense of large document collections based on topic modeling. VISTopic first extracts a set of hierarchical topics using a novel hierarchical latent tree model (HLTM) (Liu et al., 2014). In specific, a topic view accounting for the model features is designed for overall understanding and interactive exploration of the topic organization. To leverage multi-perspective information for visual analytics, VISTopic further provides an evolution view to reveal the trend of topics and a document view to show details of topical documents. Three case studies based on the dataset of IEEE VIS conference demonstrate the effectiveness of our system in gaining insights from large document collections.

Online Subset Topic Modeling For Interactive Documents Exploration

Short Text Understanding by Leveraging Knowledge into Topic Model.

Mining Coherent Topics in Documents Using Word Embeddings and Large-Scale Text Data

A Knowledge-Based Semisupervised Hierarchical Online Topic Detection Framework.

SBTM: A Joint Sentiment and Behaviour Topic Model for Online Course Discussion Forums

Interactive Topic Modeling Based on Hierarchical Dirichlet Process

TSSE-DMM: Topic Modeling for Short Texts Based on Topic Subdivision and Semantic Enhancement

Sparse online topic models

Sys-TM: A Fast and General Topic Modeling System

Short Text Topic Modeling Techniques, Applications, and Performance: A Survey

VISTopic: A visual analytics system for making sense of large document collections using hierarchical topic modeling

Modeling over Short Texts

Topic Discovery for Streaming Short Texts with CTM.

Online Visual Analytics of Text Streams

A Joint Model Of Extended Lda And Ibtm Over Streaming Chinese Short Texts

BTM: Topic Modeling over Short Texts

A Topic Model for Co-Occurring Normal Documents and Short Texts.

Interactive Visual Exploration of Topic Models using Graphs

Optimizing temporal topic segmentation for intelligent text visualization.

Automatic Text Summarization Approaches to Speed up Topic Model Learning Process

"Draw My Topics": Find Desired Topics fast from large scale of Corpus