Abstract:Event camera-based pattern recognition is a newly arising research topic in recent years. Current researchers usually transform the event streams into images, graphs, or voxels, and adopt deep neural networks for event-based classification. Although good performance can be achieved on simple event recognition datasets, however, their results may be still limited due to the following two issues. Firstly, they adopt spatial sparse event streams for recognition only, which may fail to capture the color and detailed texture information well. Secondly, they adopt either Spiking Neural Networks (SNN) for energy-efficient recognition with suboptimal results, or Artificial Neural Networks (ANN) for energy-intensive, high-performance recognition. However, seldom of them consider achieving a balance between these two aspects. In this paper, we formally propose to recognize patterns by fusing RGB frames and event streams simultaneously and propose a new RGB frame-event recognition framework to address the aforementioned issues. The proposed method contains four main modules, i.e., memory support Transformer network for RGB frame encoding, spiking neural network for raw event stream encoding, multi-modal bottleneck fusion module for RGB-Event feature aggregation, and prediction head. Due to the scarce of RGB-Event based classification dataset, we also propose a large-scale PokerEvent dataset which contains 114 classes, and 27102 frame-event pairs recorded using a DVS346 event camera. Extensive experiments on two RGB-Event based classification datasets fully validated the effectiveness of our proposed framework. We hope this work will boost the development of pattern recognition by fusing RGB frames and event streams. Both our dataset and source code of this work will be released at <a class="link-external link-https" href="https://github.com/Event-AHU/SSTFormer" rel="external noopener nofollow">this https URL</a>.

ECSNet: Spatio-Temporal Feature Learning for Event Camera

Event-based Object Detection with Lightweight Spatial Attention Mechanism

Learning SpatioTemporal and Motion Features in a Unified 2D Network for Action Recognition

E2PNet: Event to Point Cloud Registration with Spatio-Temporal Representation Learning.

Graph-based Asynchronous Event Processing for Rapid Object Recognition

Representation Learning on Event Stream via an Elastic Net-incorporated Tensor Network

Event Stream Learning Using Spatio-Temporal Event Surface

Event camera object recognition using spatiotemporal event time surface and reward-modulated spike-timing-dependent plasticity learning rule

A Lightweight Spatiotemporal Network for Online Eye Tracking with Event Camera

Spatio-temporal Focus and Lightweight Memory Network for Continuous Object Detection with Event Camera

Multi-scale Harmonic Mean Time Surfaces for Event-based Object Classification

Path-adaptive Spatio-Temporal State Space Model for Event-based Recognition with Arbitrary Duration

A Sparse Coding Multi-Scale Precise-Timing Machine Learning Algorithm for Neuromorphic Event-Based Sensors

Space-Time Event Clouds for Gesture Recognition: From RGB Cameras to Event Cameras

An Event-based Feature Representation Method for Event Stream Classification Using Deep Spiking Neural Networks

Spatiotemporal Feature Learning for Event-Based Vision

SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition

Event Voxel Set Transformer for Spatiotemporal Representation Learning on Event Streams

Asynchronous Spatio-Temporal Memory Network for Continuous Event-Based Object Detection

Continuous-time Object Segmentation using High Temporal Resolution Event Camera

Compressed Event Sensing (CES) Volumes for Event Cameras