MDMMT: Multidomain Multimodal Transformer for Video Retrieval

Maksim Dzabraev,Maksim Kalashnikov,Stepan Komkov,Aleksandr Petiushko
DOI: https://doi.org/10.1109/cvprw53098.2021.00374
2021-06-01
Abstract:We present a new state-of-the-art on the text-to-video re-trieval task on MSRVTT and LSMDC benchmarks where our model outperforms all previous solutions by a large margin. Moreover, state-of-the-art results are achieved using a single model and without finetuning. This multidomain generalisation is achieved by a proper combination of different video caption datasets. We show that our practical approach for training on different datasets can improve test results of each other. Additionally, we check intersection between many popular datasets and show that MSRVTT as well as ActivityNet contains a significant overlap between the test and the training parts. More details are available at https://github.com/papermsucode/mdmmt.
What problem does this paper attempt to address?