Abstract:Human action recognition in dark videos is a challenging task for computer vision. Recent research focuses on applying dark enhancement methods to improve the visibility of the video. However, such video processing results in the loss of critical information in the original (un-enhanced) video. Conversely, traditional two-stream methods are capable of learning information from both original and processed videos, but it can lead to a significant increase in the computational cost during the inference phase in the task of video classification. To address these challenges, we propose a novel teacher-student video classification framework, named Dual-Light KnowleDge Distillation for Action Recognition in the Dark (DL-KDD). This framework enables the model to learn from both original and enhanced video without introducing additional computational cost during inference. Specifically, DL-KDD utilizes the strategy of knowledge distillation during training. The teacher model is trained with enhanced video, and the student model is trained with both the original video and the soft target generated by the teacher model. This teacher-student framework allows the student model to predict action using only the original input video during inference. In our experiments, the proposed DL-KDD framework outperforms state-of-the-art methods on the ARID, ARID V1.5, and Dark-48 datasets. We achieve the best performance on each dataset and up to a 4.18% improvement on Dark-48, using only original video inputs, thus avoiding the use of two-stream framework or enhancement modules for inference. We further validate the effectiveness of the distillation strategy in ablative experiments. The results highlight the advantages of our knowledge distillation framework in dark human action recognition.

TKD: Temporal Knowledge Distillation for Active Perception

Self-Paced Knowledge Distillation for Real-Time Image Guided Depth Completion

Temporal Distinct Representation Learning for Action Recognition

Distilling Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection

Temporal Dynamic Graph LSTM for Action-driven Video Object Detection

Delta Distillation for Efficient Video Processing

TOD: Transprecise Object Detection to Maximise Real-Time Accuracy on the Edge

TLS-RWKV: Real-Time Online Action Detection with Temporal Label Smoothing

Towards Efficient 3D Object Detection with Knowledge Distillation

TDN: Temporal Difference Networks for Efficient Action Recognition

Collaborative spatial-temporal distillation for efficient video deraining

DL-KDD: Dual-Light Knowledge Distillation for Action Recognition in the Dark

Deep Learning and Hybrid Approaches for Dynamic Scene Analysis, Object Detection and Motion Tracking

Perception Over Time: Temporal Dynamics for Robust Image Understanding

Temporally Identity-Aware SSD With Attentional LSTM

TKN: Transformer-based Keypoint Prediction Network For Real-time Video Prediction

ODTrack: Online Dense Temporal Token Learning for Visual Tracking

TEDdet: Temporal Feature Exchange and Difference Network for Online Real-Time Action Detection

Structural Knowledge Distillation for Object Detection

Dynamic Knowledge Distillation with Noise Elimination for RGB-D Salient Object Detection

Cyclic Refiner: Object-Aware Temporal Representation Learning for Multi-view 3D Detection and Tracking