Hierarchical Variational Network for User-Diversified & Query-Focused Video Summarization.

Pin Jiang,Yahong Han
DOI: https://doi.org/10.1145/3323873.3325040
2019-01-01
Abstract:This paper focuses on the query-focused video summarization, which is an extended task of video summarization and aims to automatically generate user-oriented summary by highlighting frames/shots relevant to the query. This task is different from traditional video summarization in paying attention to users' subjectivity through queries. Diversity is a recognized important property in video summarization. However, existing methods only consider diversity as the dissimilarity between frames/shots which is far from user-oriented summarization. Users' different understandings of video should be an important source of diversity, reflected in the process of eliminating query-unrelated redundancy. To this end, this paper explores user-diversified & query-focused video summarization via a well-devised hierarchical variational network called HVN. HVN has three distinctive characteristics: (i) it has a hierarchical structure to model query-related long-range temporal dependency; (ii) it employs diverse attention mechanisms to encode query-related and context-important information and makes them balanced; (iii) it employs a multilevel self-attention module and a variational autoencoder module to add user-oriented diversity and stochastic factors. Experimental results demonstrate that HVN not only outperforms the state-of-the-arts but also improves the user-oriented diversity to some extent.
What problem does this paper attempt to address?