OmDet: Large‐scale vision‐language multi‐dataset pre‐training with multimodal detection network

Tiancheng Zhao,Peng Liu,Kyusong Lee
DOI: https://doi.org/10.1049/cvi2.12268
IF: 1.484
2024-01-26
IET Computer Vision
Abstract:OmDet, a novel language‐aware detector, designed to enhance open‐vocabulary and open‐world object detection through a continual learning approach and multi‐dataset vision‐language pre‐training is presented. By using natural language for knowledge representation, the authors successfully increase the "visual vocabularyˮ; and create a unified, language‐conditioned detection framework, which outperforms previous models on object detection and phrase grounding. This promising method proves the effectiveness of joint learning from multiple datasets and presents a path forward for scaling to even larger datasets. The advancement of object detection (OD) in open‐vocabulary and open‐world scenarios is a critical challenge in computer vision. OmDet, a novel language‐aware object detection architecture and an innovative training mechanism that harnesses continual learning and multi‐dataset vision‐language pre‐training is introduced. Leveraging natural language as a universal knowledge representation, OmDet accumulates "visual vocabularies" from diverse datasets, unifying the task as a language‐conditioned detection framework. The multimodal detection network (MDN) overcomes the challenges of multi‐dataset joint training and generalizes to numerous training datasets without manual label taxonomy merging. The authors demonstrate superior performance of OmDet over strong baselines in object detection in the wild, open‐vocabulary detection, and phrase grounding, achieving state‐of‐the‐art results. Ablation studies reveal the impact of scaling the pre‐training visual vocabulary, indicating a promising direction for further expansion to larger datasets. The effectiveness of our deep fusion approach is underscored by its ability to learn jointly from multiple datasets, enhancing performance through knowledge sharing.
computer science, artificial intelligence,engineering, electrical & electronic
What problem does this paper attempt to address?