Abstract:We initiate the first empirical study on the use of MLP architectures for vision-and-language (VL) fusion. Through extensive experiments on 5 VL tasks and 5 robust VQA benchmarks, we find that: (i) Without pre-training, using MLPs for multimodal fusion has a noticeable performance gap compared to transformers; (ii) However, VL pre-training can help close the performance gap; (iii) Instead of heavy multi-head attention, adding tiny one-head attention to MLPs is sufficient to achieve comparable performance to transformers. Moreover, we also find that the performance gap between MLPs and transformers is not widened when being evaluated on the harder robust VQA benchmarks, suggesting using MLPs for VL fusion can generalize roughly to a similar degree as using transformers. These results hint that MLPs can effectively learn to align vision and text features extracted from lower-level encoders without heavy reliance on self-attention. Based on this, we ask an even bolder question: can we have an all-MLP architecture for VL modeling, where both VL fusion and the vision encoder are replaced with MLPs? Our result shows that an all-MLP VL model is sub-optimal compared to state-of-the-art full-featured VL models when both of them get pre-trained. However, pre-training an all-MLP can surprisingly achieve a better average score than full-featured transformer models without pre-training. This indicates the potential of large-scale pre-training of MLP-like architectures for VL modeling and inspires the future research direction on simplifying well-established VL modeling with less inductive design bias. Our code is publicly available at: <a class="link-external link-https" href="https://github.com/easonnie/mlp-vil" rel="external noopener nofollow">this https URL</a>

Strategies for using MLP based features with limited target-language training data.

Cross-Lingual and Ensemble MLPs Strategies for Low-Resource Speech Recognition

Articulatory Feature Based Multilingual MLPs for Low-Resource Speech Recognition.

AudioVSR: Enhancing Video Speech Recognition with Audio Data

Multi-Stream Posterior Features and Combining Subspace Gmms for Low Resource Lvcsr

VarASV: Enabling Pitch-variable Automatic Speaker Verification Via Multi-task Learning

LEARNING CROSS-LINGUAL INFORMATION WITH MULTILINGUAL BLSTM FOR SPEECH SYNTHESIS OF LOW-RESOURCE LANGUAGES

Combination Of Data Borrowing Strategies For Low-Resource Lvcsr

MLP-HMM Two-Stage Unsupervised Training for Low-Resource Languages on Conversational Telephone Speech Recognition

Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data

Accent Recognition with Hybrid Phonetic Features

ON MODULAR TRAINING OF NEURAL ACOUSTICS-TO-WORD MODEL FOR LVCSR

Improvement of Acoustic Models Fused with Lip Visual Information for Low-Resource Speech

Enhancing multilingual speech recognition in air traffic control by sentence-level language identification

Leveraging native language information for improved accented speech recognition

MLP Architectures for Vision-and-Language Modeling: An Empirical Study

Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition

Auxiliary Features from Laser-Doppler Vibrometer Sensor for Deep Neural Network Based Robust Speech Recognition

A multilingual training strategy for low resource Text to Speech

Towards High Performance LVCSR in Speech-to-Speech Translation System on Smart Phones.

Investigating the Impact of Cross-lingual Acoustic-Phonetic Similarities on Multilingual Speech Recognition