CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Rujia, Gao, Xiangbo, Xiang, Hao, Xu, Runsheng, Tu, Zhengzhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STAMP: Scalable Task And Model-agnostic Collaborative Perception
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
LangCoop: Collaborative Driving with Language
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
by: Godbole, Mihir, et al.
Published: (2025)
by: Godbole, Mihir, et al.
Published: (2025)
Background Fades, Foreground Leads: Curriculum-Guided Background Pruning for Efficient Foreground-Centric Collaborative Perception
by: Wu, Yuheng, et al.
Published: (2025)
by: Wu, Yuheng, et al.
Published: (2025)
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
Automated Vehicles Should be Connected with Natural Language
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023)
by: Man, Yunze, et al.
Published: (2023)
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
by: Polizzi, Vincenzo, et al.
Published: (2026)
by: Polizzi, Vincenzo, et al.
Published: (2026)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving
by: Xu, Dongyang, et al.
Published: (2024)
by: Xu, Dongyang, et al.
Published: (2024)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation
by: Zhuang, Lipeng, et al.
Published: (2024)
by: Zhuang, Lipeng, et al.
Published: (2024)
ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
by: Xie, Weidong, et al.
Published: (2024)
by: Xie, Weidong, et al.
Published: (2024)
On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
by: Wu, Yiming, et al.
Published: (2025)
by: Wu, Yiming, et al.
Published: (2025)
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
by: Han, Yu, et al.
Published: (2025)
by: Han, Yu, et al.
Published: (2025)
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
by: Yu, Jiongze, et al.
Published: (2026)
by: Yu, Jiongze, et al.
Published: (2026)
PISCO: Precise Video Instance Insertion with Sparse Control
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
by: Chuang, Ian, et al.
Published: (2025)
by: Chuang, Ian, et al.
Published: (2025)
Efficient Perception, Planning, and Control Algorithm for Vision-Based Automated Vehicles
by: Lee, Der-Hau
Published: (2022)
by: Lee, Der-Hau
Published: (2022)
Interruption-Aware Cooperative Perception for V2X Communication-Aided Autonomous Driving
by: Ren, Shunli, et al.
Published: (2023)
by: Ren, Shunli, et al.
Published: (2023)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
Online,Target-Free LiDAR-Camera Extrinsic Calibration via Cross-Modal Mask Matching
by: Huang, Zhiwei, et al.
Published: (2024)
by: Huang, Zhiwei, et al.
Published: (2024)
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
by: Guo, Jianing, et al.
Published: (2025)
by: Guo, Jianing, et al.
Published: (2025)
Region-R1: Reinforcing Query-Side Region Cropping for Multi-Modal Re-Ranking
by: Hu, Chan-Wei, et al.
Published: (2026)
by: Hu, Chan-Wei, et al.
Published: (2026)
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025)
by: Barron, William, et al.
Published: (2025)
CageDroneRF: A Large-Scale RF Benchmark and Toolkit for Drone Perception
by: Rostami, Mohammad, et al.
Published: (2026)
by: Rostami, Mohammad, et al.
Published: (2026)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
by: Xu, Huaiyuan, et al.
Published: (2024)
by: Xu, Huaiyuan, et al.
Published: (2024)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
by: Xia, Shao-Jun, et al.
Published: (2025)
by: Xia, Shao-Jun, et al.
Published: (2025)
Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model
by: Zhang, Ruibin, et al.
Published: (2024)
by: Zhang, Ruibin, et al.
Published: (2024)
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
by: Liu, Minghuan, et al.
Published: (2025)
by: Liu, Minghuan, et al.
Published: (2025)
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
by: Wen, Jiawen, et al.
Published: (2026)
by: Wen, Jiawen, et al.
Published: (2026)
SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2025)
by: Zhang, Haiming, et al.
Published: (2025)
Learning to Navigate Socially Through Proactive Risk Perception
by: Xiao, Erjia, et al.
Published: (2025)
by: Xiao, Erjia, et al.
Published: (2025)
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
by: Tang, Zecong, et al.
Published: (2026)
by: Tang, Zecong, et al.
Published: (2026)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
Similar Items
-
STAMP: Scalable Task And Model-agnostic Collaborative Perception
by: Gao, Xiangbo, et al.
Published: (2025) -
LangCoop: Collaborative Driving with Language
by: Gao, Xiangbo, et al.
Published: (2025) -
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
by: Godbole, Mihir, et al.
Published: (2025) -
Background Fades, Foreground Leads: Curriculum-Guided Background Pruning for Efficient Foreground-Centric Collaborative Perception
by: Wu, Yuheng, et al.
Published: (2025) -
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
by: Gao, Xiangbo, et al.
Published: (2025)