Guest Editorial Introduction to the Special Section on Intelligent Visual Content Analysis and Understanding

Hongliang Li,Lu Fang,Tianzhu Zhang
DOI: https://doi.org/10.1109/tcsvt.2020.3031416
2020-01-01
Abstract:Visual content analysis and understanding attract tremendous attention because of its potentially wide range of applications including human activity analysis, automated photo face tagging, multicamera tracking, crowded counting, and biometric security. With recent progress in end-to-end differentiable learning, the accuracy of algorithms has been significantly improved and even outperforms humans in some tasks. In addition, multimodality methods, targeting on making full use of various visual data sources, are further investigated. These developments contribute to the innovations of two core modules for a typical intelligent vision system, i.e., image and video description and recognition, which are critical for the success of the visual content analysis and understanding in more complex and challenging open world.
What problem does this paper attempt to address?