CC-LSTM: Cross and Conditional Long-Short Time Memory for Video Captioning

Jiangbo Ai,Yang Yang,Xing Xu,Jie Zhou,Heng Tao Shen
DOI: https://doi.org/10.1007/978-3-030-68780-9_30
2020-01-01
Abstract:Automatically generating natural language descriptions for in-the-wild videos is a challenging task. Most recent progress in this field has been made through the combination of Convolutional Neural Networks (CNNs) and Encoder-Decoder Recurrent Neural Networks (RNNs). However, existing Encoder-Decoder RNNs framework has difficulty in capturing a large number of long-range dependencies along with the increasing of the number of LSTM units. It brings a vast information loss and leads to poor performance for our task. To explore this problem, in this paper, we propose a novel framework, namely Cross and Conditional Long Short-Term Memory (CC-LSTM). It is composed of a novel Cross Long Short-Term Memory (Cr-LSTM) for the encoding module and Conditional Long Short-Term Memory (Co-LSTM) for the decoding module. In the encoding module, the Cr-LSTM encodes the visual input into a richly informative representation by a cross-input method. In the decoding module, the Co-LSTM feeds the visual features, which is based on generated sentence and contains the global information of the visual content, into the LSTM unit as an extra visual feature. For the work of video capturing, extensive experiments are conducted on two public datasets, i.e., MSVD and MSR-VTT. Along with visualizing the results and how our model works, these experiments quantitatively demonstrate the effectiveness of the proposed CC-LSTM on translating videos to sentences with rich semantics.
What problem does this paper attempt to address?