SiMO: Single-Modality-Operable Multimodal Collaborative Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Jiageng, Zhao, Shengjie, Li, Bing, Huang, Jiafeng, Ye, Kenan, Deng, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linking Modality Isolation in Heterogeneous Collaborative Perception
by: Liu, Changxing, et al.
Published: (2026)
by: Liu, Changxing, et al.
Published: (2026)
V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR Localization
by: Lin, Wenkai, et al.
Published: (2025)
by: Lin, Wenkai, et al.
Published: (2025)
Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
by: Zhu, Zhihao, et al.
Published: (2026)
by: Zhu, Zhihao, et al.
Published: (2026)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
by: Deria, Ankan, et al.
Published: (2026)
by: Deria, Ankan, et al.
Published: (2026)
CHARM: Collaborative Harmonization across Arbitrary Modalities for Modality-agnostic Semantic Segmentation
by: Wen, Lekang, et al.
Published: (2025)
by: Wen, Lekang, et al.
Published: (2025)
Modality Invariant Multimodal Learning to Handle Missing Modalities: A Single-Branch Approach
by: Saeed, Muhammad Saad, et al.
Published: (2024)
by: Saeed, Muhammad Saad, et al.
Published: (2024)
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
by: Liang, Yongyuan, et al.
Published: (2025)
by: Liang, Yongyuan, et al.
Published: (2025)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
by: Li, Teng, et al.
Published: (2025)
by: Li, Teng, et al.
Published: (2025)
Deep Feature Gaussian Processes for Single-Scene Aerosol Optical Depth Reconstruction
by: Liu, Shengjie, et al.
Published: (2024)
by: Liu, Shengjie, et al.
Published: (2024)
Learning from Rendering: Realistic and Controllable Extreme Rainy Image Synthesis for Autonomous Driving Simulation
by: Zhou, Kaibin, et al.
Published: (2025)
by: Zhou, Kaibin, et al.
Published: (2025)
ActFormer: Scalable Collaborative Perception via Active Queries
by: Huang, Suozhi, et al.
Published: (2024)
by: Huang, Suozhi, et al.
Published: (2024)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
by: Hao, Haihong, et al.
Published: (2025)
by: Hao, Haihong, et al.
Published: (2025)
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception
by: Li, Xilai, et al.
Published: (2025)
by: Li, Xilai, et al.
Published: (2025)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024)
by: Leng, Sicong, et al.
Published: (2024)
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
by: Wu, Jiahe, et al.
Published: (2026)
by: Wu, Jiahe, et al.
Published: (2026)
GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction
by: Huang, Yi-Chuan, et al.
Published: (2025)
by: Huang, Yi-Chuan, et al.
Published: (2025)
You Share Beliefs, I Adapt: Progressive Heterogeneous Collaborative Perception
by: Si, Hao, et al.
Published: (2025)
by: Si, Hao, et al.
Published: (2025)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
by: Guo, Xiangyu, et al.
Published: (2025)
by: Guo, Xiangyu, et al.
Published: (2025)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
by: Li, Yuyang, et al.
Published: (2025)
by: Li, Yuyang, et al.
Published: (2025)
Spotlight on Token Perception for Multimodal Reinforcement Learning
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
Multimodal LLMs under Pairwise Modalities
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
End-to-End 3D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
by: Yang, Zhenwei, et al.
Published: (2025)
by: Yang, Zhenwei, et al.
Published: (2025)
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
by: Wang, Rujia, et al.
Published: (2025)
by: Wang, Rujia, et al.
Published: (2025)
Is Discretization Fusion All You Need for Collaborative Perception?
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
by: Liu, Ruiheng, et al.
Published: (2026)
by: Liu, Ruiheng, et al.
Published: (2026)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
Self-Localized Collaborative Perception
by: Ni, Zhenyang, et al.
Published: (2024)
by: Ni, Zhenyang, et al.
Published: (2024)
SparseFusion: Efficient Sparse Multi-Modal Fusion Framework for Long-Range 3D Perception
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
Evolving Without Ending: Unifying Multimodal Incremental Learning for Continual Panoptic Perception
by: Yuan, Bo, et al.
Published: (2026)
by: Yuan, Bo, et al.
Published: (2026)
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
by: Xu, Hengbo, et al.
Published: (2026)
by: Xu, Hengbo, et al.
Published: (2026)
U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
FeaKM: Robust Collaborative Perception under Noisy Pose Conditions
by: Hao, Jiuwu, et al.
Published: (2025)
by: Hao, Jiuwu, et al.
Published: (2025)
Collaborative Perception Datasets for Autonomous Driving: A Review
by: Wang, Naibang, et al.
Published: (2025)
by: Wang, Naibang, et al.
Published: (2025)
PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction
by: Hou, Xiaolu, et al.
Published: (2025)
by: Hou, Xiaolu, et al.
Published: (2025)
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Similar Items
-
Linking Modality Isolation in Heterogeneous Collaborative Perception
by: Liu, Changxing, et al.
Published: (2026) -
V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR Localization
by: Lin, Wenkai, et al.
Published: (2025) -
Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs
by: Chen, Junhao, et al.
Published: (2024) -
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
by: Zhu, Zhihao, et al.
Published: (2026) -
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)