CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Wentao, Wang, Xiao, Li, Chenglong, Jiang, Bo, Tang, Jin, Luo, Bin, Liu, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
CRSOT: Cross-Resolution Object Tracking using Unaligned Frame and Event Cameras
by: Zhu, Yabin, et al.
Published: (2024)
by: Zhu, Yabin, et al.
Published: (2024)
Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction
by: Jiang, Jianping, et al.
Published: (2024)
by: Jiang, Jianping, et al.
Published: (2024)
Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
EventDiff: A Unified and Efficient Diffusion Model Framework for Event-based Video Frame Interpolation
by: Zheng, Hanle, et al.
Published: (2025)
by: Zheng, Hanle, et al.
Published: (2025)
Mamba-FETrack: Frame-Event Tracking via State Space Model
by: Huang, Ju, et al.
Published: (2024)
by: Huang, Ju, et al.
Published: (2024)
SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
by: Liu, Jiaxiong, et al.
Published: (2024)
by: Liu, Jiaxiong, et al.
Published: (2024)
RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
A Unified Framework for Event-based Frame Interpolation with Ad-hoc Deblurring in the Wild
by: Sun, Lei, et al.
Published: (2023)
by: Sun, Lei, et al.
Published: (2023)
Bimodal SegNet: Instance Segmentation Fusing Events and RGB Frames for Robotic Grasping
by: Kachole, Sanket, et al.
Published: (2023)
by: Kachole, Sanket, et al.
Published: (2023)
Unified Arbitrary-Time Video Frame Interpolation and Prediction
by: Jin, Xin, et al.
Published: (2025)
by: Jin, Xin, et al.
Published: (2025)
LiFR-Seg: Anytime High-Frame-Rate Segmentation via Event-Guided Propagation
by: Wu, Xiaoshan, et al.
Published: (2026)
by: Wu, Xiaoshan, et al.
Published: (2026)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
Unleashing the Power of CNN and Transformer for Balanced RGB-Event Video Recognition
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models
by: Xu, Gang, et al.
Published: (2026)
by: Xu, Gang, et al.
Published: (2026)
Deep Visual Odometry with Events and Frames
by: Pellerito, Roberto, et al.
Published: (2023)
by: Pellerito, Roberto, et al.
Published: (2023)
EvRainDrop: HyperGraph-guided Completion for Effective Frame and Event Stream Aggregation
by: Wang, Futian, et al.
Published: (2025)
by: Wang, Futian, et al.
Published: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
by: Cao, Meiqi, et al.
Published: (2024)
by: Cao, Meiqi, et al.
Published: (2024)
Adversarial Attack for RGB-Event based Visual Object Tracking
by: Chen, Qiang, et al.
Published: (2025)
by: Chen, Qiang, et al.
Published: (2025)
BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation using RGB Frames and Events
by: Li, Yijin, et al.
Published: (2024)
by: Li, Yijin, et al.
Published: (2024)
Pre-training on High Definition X-ray Images: An Experimental Study
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Tinted Frames: Question Framing Blinds Vision-Language Models
by: Fan, Wan-Cyuan, et al.
Published: (2026)
by: Fan, Wan-Cyuan, et al.
Published: (2026)
ADV2E: Bridging the Gap Between Analogue Circuit and Discrete Frames in the Video-to-Events Simulator
by: Jiang, Xiao, et al.
Published: (2024)
by: Jiang, Xiao, et al.
Published: (2024)
Unifying Graph Contrastive Learning via Graph Message Augmentation
by: Zhang, Ziyan, et al.
Published: (2024)
by: Zhang, Ziyan, et al.
Published: (2024)
Self-supervised Learning of Event-guided Video Frame Interpolation for Rolling Shutter Frames
by: Lu, Yunfan, et al.
Published: (2023)
by: Lu, Yunfan, et al.
Published: (2023)
Event-based Continuous Color Video Decompression from Single Frames
by: Wang, Ziyun, et al.
Published: (2023)
by: Wang, Ziyun, et al.
Published: (2023)
Fully Spiking Neural Networks for Unified Frame-Event Object Tracking
by: Yang, Jingjun, et al.
Published: (2025)
by: Yang, Jingjun, et al.
Published: (2025)
RGB-Event HyperGraph Prompt for Kilometer Marker Recognition based on Pre-trained Foundation Models
by: Xian, Xiaoyu, et al.
Published: (2026)
by: Xian, Xiaoyu, et al.
Published: (2026)
F2M-Reg: Unsupervised RGB-D Point Cloud Registration with Frame-to-Model Optimization
by: Yu, Zhinan, et al.
Published: (2024)
by: Yu, Zhinan, et al.
Published: (2024)
LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning
by: Chao, Lianying, et al.
Published: (2026)
by: Chao, Lianying, et al.
Published: (2026)
Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach
by: Li, Hebei, et al.
Published: (2025)
by: Li, Hebei, et al.
Published: (2025)
Raw2Event: Converting Raw Frame Camera into Event Camera
by: Ning, Zijie, et al.
Published: (2025)
by: Ning, Zijie, et al.
Published: (2025)
Hybrid Event Frame Sensors: Modeling, Calibration, and Simulation
by: Lu, Yunfan, et al.
Published: (2025)
by: Lu, Yunfan, et al.
Published: (2025)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
Event-Based Video Frame Interpolation With Cross-Modal Asymmetric Bidirectional Motion Fields
by: Kim, Taewoo, et al.
Published: (2025)
by: Kim, Taewoo, et al.
Published: (2025)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
by: Yu, Sicheng, et al.
Published: (2024)
by: Yu, Sicheng, et al.
Published: (2024)
Similar Items
-
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025) -
CRSOT: Cross-Resolution Object Tracking using Unaligned Frame and Event Cameras
by: Zhu, Yabin, et al.
Published: (2024) -
Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction
by: Jiang, Jianping, et al.
Published: (2024) -
Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms
by: Wang, Xiao, et al.
Published: (2025) -
EventDiff: A Unified and Efficient Diffusion Model Framework for Event-based Video Frame Interpolation
by: Zheng, Hanle, et al.
Published: (2025)