Part-Whole Relational Fusion Towards Multi-Modal Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yi, Li, Chengxin, Xu, Shoukun, Han, Jungong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
Mamba Capsule Routing Towards Part-Whole Relational Camouflaged Object Detection
von: Zhang, Dingwen, et al.
Veröffentlicht: (2024)
von: Zhang, Dingwen, et al.
Veröffentlicht: (2024)
SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality
von: Li, Sijie, et al.
Veröffentlicht: (2025)
von: Li, Sijie, et al.
Veröffentlicht: (2025)
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts
von: Zhu, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhu, Chenyang, et al.
Veröffentlicht: (2024)
Modality Prompts for Arbitrary Modality Salient Object Detection
von: Huang, Nianchang, et al.
Veröffentlicht: (2024)
von: Huang, Nianchang, et al.
Veröffentlicht: (2024)
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
von: Shi, Yi, et al.
Veröffentlicht: (2025)
von: Shi, Yi, et al.
Veröffentlicht: (2025)
Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
von: Liu, Weide, et al.
Veröffentlicht: (2025)
von: Liu, Weide, et al.
Veröffentlicht: (2025)
Robust Multi-Modal Image Stitching for Improved Scene Understanding
von: Dutta, Aritra, et al.
Veröffentlicht: (2023)
von: Dutta, Aritra, et al.
Veröffentlicht: (2023)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy Prediction
von: Lv, Chengxin, et al.
Veröffentlicht: (2026)
von: Lv, Chengxin, et al.
Veröffentlicht: (2026)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
Towards Online Multi-Modal Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
von: Jiang, Hanqi, et al.
Veröffentlicht: (2024)
von: Jiang, Hanqi, et al.
Veröffentlicht: (2024)
RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
von: Song, Ziying, et al.
Veröffentlicht: (2024)
von: Song, Ziying, et al.
Veröffentlicht: (2024)
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
von: Chen, Jingkun, et al.
Veröffentlicht: (2026)
von: Chen, Jingkun, et al.
Veröffentlicht: (2026)
InfScene-SR: Arbitrary-Size Image Super-Resolution via Iterative Joint-Denoising
von: Sun, Shoukun, et al.
Veröffentlicht: (2026)
von: Sun, Shoukun, et al.
Veröffentlicht: (2026)
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
von: Hou, Minghui, et al.
Veröffentlicht: (2025)
Salient Object Detection From Arbitrary Modalities
von: Huang, Nianchang, et al.
Veröffentlicht: (2024)
von: Huang, Nianchang, et al.
Veröffentlicht: (2024)
Micro-macro Gaussian Splatting with Enhanced Scalability for Unconstrained Scene Reconstruction
von: Li, Yihui, et al.
Veröffentlicht: (2025)
von: Li, Yihui, et al.
Veröffentlicht: (2025)
Towards Efficient Information Fusion: Concentric Dual Fusion Attention Based Multiple Instance Learning for Whole Slide Images
von: Liu, Yujian, et al.
Veröffentlicht: (2024)
von: Liu, Yujian, et al.
Veröffentlicht: (2024)
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
von: Yin, Xingyilang, et al.
Veröffentlicht: (2025)
von: Yin, Xingyilang, et al.
Veröffentlicht: (2025)
Equivariant Multi-Modality Image Fusion
von: Zhao, Zixiang, et al.
Veröffentlicht: (2023)
von: Zhao, Zixiang, et al.
Veröffentlicht: (2023)
Guided and Variance-Corrected Fusion with One-shot Style Alignment for Large-Content Image Generation
von: Sun, Shoukun, et al.
Veröffentlicht: (2024)
von: Sun, Shoukun, et al.
Veröffentlicht: (2024)
Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
Tracking and Segmenting Anything in Any Modality
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
Relation-Aware Meta-Learning for Zero-shot Sketch-Based Image Retrieval
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios
von: Xu, Jialei, et al.
Veröffentlicht: (2024)
von: Xu, Jialei, et al.
Veröffentlicht: (2024)
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
von: Rohrmoser, Nikolo, et al.
Veröffentlicht: (2026)
von: Rohrmoser, Nikolo, et al.
Veröffentlicht: (2026)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
FusionSAM: Visual Multi-Modal Learning with Segment Anything
von: Li, Daixun, et al.
Veröffentlicht: (2024)
von: Li, Daixun, et al.
Veröffentlicht: (2024)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
von: Zheng, Jinzhi, et al.
Veröffentlicht: (2024)
von: Zheng, Jinzhi, et al.
Veröffentlicht: (2024)
RepViT-SAM: Towards Real-Time Segmenting Anything
von: Wang, Ao, et al.
Veröffentlicht: (2023)
von: Wang, Ao, et al.
Veröffentlicht: (2023)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
von: Ding, Xinpeng, et al.
Veröffentlicht: (2024)
von: Ding, Xinpeng, et al.
Veröffentlicht: (2024)
Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes
von: Agarwal, Nakul, et al.
Veröffentlicht: (2026)
von: Agarwal, Nakul, et al.
Veröffentlicht: (2026)
On Exploring PDE Modeling for Point Cloud Video Representation Learning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2024)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2024)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
von: Chen, Anjun, et al.
Veröffentlicht: (2024)
von: Chen, Anjun, et al.
Veröffentlicht: (2024)
IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans Fusion
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
von: Liu, Yi, et al.
Veröffentlicht: (2024) -
Mamba Capsule Routing Towards Part-Whole Relational Camouflaged Object Detection
von: Zhang, Dingwen, et al.
Veröffentlicht: (2024) -
SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality
von: Li, Sijie, et al.
Veröffentlicht: (2025) -
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts
von: Zhu, Chenyang, et al.
Veröffentlicht: (2024) -
Modality Prompts for Arbitrary Modality Salient Object Detection
von: Huang, Nianchang, et al.
Veröffentlicht: (2024)