Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yi, Li, Chengxin, Xu, Shoukun, Han, Jungong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Mamba Capsule Routing Towards Part-Whole Relational Camouflaged Object Detection
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality
by: Li, Sijie, et al.
Published: (2025)
by: Li, Sijie, et al.
Published: (2025)
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts
by: Zhu, Chenyang, et al.
Published: (2024)
by: Zhu, Chenyang, et al.
Published: (2024)
Modality Prompts for Arbitrary Modality Salient Object Detection
by: Huang, Nianchang, et al.
Published: (2024)
by: Huang, Nianchang, et al.
Published: (2024)
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
by: Shi, Yi, et al.
Published: (2025)
by: Shi, Yi, et al.
Published: (2025)
Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
by: Liu, Weide, et al.
Published: (2025)
by: Liu, Weide, et al.
Published: (2025)
Robust Multi-Modal Image Stitching for Improved Scene Understanding
by: Dutta, Aritra, et al.
Published: (2023)
by: Dutta, Aritra, et al.
Published: (2023)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy Prediction
by: Lv, Chengxin, et al.
Published: (2026)
by: Lv, Chengxin, et al.
Published: (2026)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Towards Online Multi-Modal Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2025)
by: Li, Xinpeng, et al.
Published: (2025)
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
by: Jiang, Hanqi, et al.
Published: (2024)
by: Jiang, Hanqi, et al.
Published: (2024)
RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
by: Chen, Jingkun, et al.
Published: (2026)
by: Chen, Jingkun, et al.
Published: (2026)
InfScene-SR: Arbitrary-Size Image Super-Resolution via Iterative Joint-Denoising
by: Sun, Shoukun, et al.
Published: (2026)
by: Sun, Shoukun, et al.
Published: (2026)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
by: Hou, Minghui, et al.
Published: (2025)
by: Hou, Minghui, et al.
Published: (2025)
Salient Object Detection From Arbitrary Modalities
by: Huang, Nianchang, et al.
Published: (2024)
by: Huang, Nianchang, et al.
Published: (2024)
Micro-macro Gaussian Splatting with Enhanced Scalability for Unconstrained Scene Reconstruction
by: Li, Yihui, et al.
Published: (2025)
by: Li, Yihui, et al.
Published: (2025)
Towards Efficient Information Fusion: Concentric Dual Fusion Attention Based Multiple Instance Learning for Whole Slide Images
by: Liu, Yujian, et al.
Published: (2024)
by: Liu, Yujian, et al.
Published: (2024)
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
by: Yin, Xingyilang, et al.
Published: (2025)
by: Yin, Xingyilang, et al.
Published: (2025)
Equivariant Multi-Modality Image Fusion
by: Zhao, Zixiang, et al.
Published: (2023)
by: Zhao, Zixiang, et al.
Published: (2023)
Guided and Variance-Corrected Fusion with One-shot Style Alignment for Large-Content Image Generation
by: Sun, Shoukun, et al.
Published: (2024)
by: Sun, Shoukun, et al.
Published: (2024)
Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track
by: Gao, Mingqi, et al.
Published: (2026)
by: Gao, Mingqi, et al.
Published: (2026)
Tracking and Segmenting Anything in Any Modality
by: Zhang, Tianlu, et al.
Published: (2025)
by: Zhang, Tianlu, et al.
Published: (2025)
Relation-Aware Meta-Learning for Zero-shot Sketch-Based Image Retrieval
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios
by: Xu, Jialei, et al.
Published: (2024)
by: Xu, Jialei, et al.
Published: (2024)
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
by: Rohrmoser, Nikolo, et al.
Published: (2026)
by: Rohrmoser, Nikolo, et al.
Published: (2026)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
FusionSAM: Visual Multi-Modal Learning with Segment Anything
by: Li, Daixun, et al.
Published: (2024)
by: Li, Daixun, et al.
Published: (2024)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
by: Zheng, Jinzhi, et al.
Published: (2024)
by: Zheng, Jinzhi, et al.
Published: (2024)
RepViT-SAM: Towards Real-Time Segmenting Anything
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
by: Ding, Xinpeng, et al.
Published: (2024)
by: Ding, Xinpeng, et al.
Published: (2024)
Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes
by: Agarwal, Nakul, et al.
Published: (2026)
by: Agarwal, Nakul, et al.
Published: (2026)
On Exploring PDE Modeling for Point Cloud Video Representation Learning
by: Huang, Zhuoxu, et al.
Published: (2024)
by: Huang, Zhuoxu, et al.
Published: (2024)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
by: Tan, Jiawei, et al.
Published: (2024)
by: Tan, Jiawei, et al.
Published: (2024)
AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
by: Chen, Anjun, et al.
Published: (2024)
by: Chen, Anjun, et al.
Published: (2024)
IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
Similar Items
-
Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
by: Liu, Yi, et al.
Published: (2024) -
Mamba Capsule Routing Towards Part-Whole Relational Camouflaged Object Detection
by: Zhang, Dingwen, et al.
Published: (2024) -
SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality
by: Li, Sijie, et al.
Published: (2025) -
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts
by: Zhu, Chenyang, et al.
Published: (2024) -
Modality Prompts for Arbitrary Modality Salient Object Detection
by: Huang, Nianchang, et al.
Published: (2024)