Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Zehui, Zhang, Chenqi, Wang, Mingru, Wei, Minuo, Cheng, Shiwei, Guan, Cuntai, Han, Ting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Explaining Multi-modal Large Language Models by Analyzing their Vision Perception
di: Giulivi, Loris, et al.
Pubblicazione: (2024)
di: Giulivi, Loris, et al.
Pubblicazione: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
di: Zhang, Ming, et al.
Pubblicazione: (2024)
di: Zhang, Ming, et al.
Pubblicazione: (2024)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
di: Miao, Yongzhu, et al.
Pubblicazione: (2023)
di: Miao, Yongzhu, et al.
Pubblicazione: (2023)
APEX: Learning Adaptive Priorities for Multi-Objective Alignment in Vision-Language Generation
di: Chen, Dongliang, et al.
Pubblicazione: (2026)
di: Chen, Dongliang, et al.
Pubblicazione: (2026)
Hierarchical Vision-Language Interaction for Facial Action Unit Detection
di: Li, Yong, et al.
Pubblicazione: (2026)
di: Li, Yong, et al.
Pubblicazione: (2026)
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
di: Kim, Hayeon, et al.
Pubblicazione: (2026)
di: Kim, Hayeon, et al.
Pubblicazione: (2026)
MindFormer: Semantic Alignment of Multi-Subject fMRI for Brain Decoding
di: Han, Inhwa, et al.
Pubblicazione: (2024)
di: Han, Inhwa, et al.
Pubblicazione: (2024)
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
di: Zhu, Bin, et al.
Pubblicazione: (2023)
di: Zhu, Bin, et al.
Pubblicazione: (2023)
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
di: Zhang, Xinyao, et al.
Pubblicazione: (2026)
di: Zhang, Xinyao, et al.
Pubblicazione: (2026)
MedMAP: Promoting Incomplete Multi-modal Brain Tumor Segmentation with Alignment
di: Liu, Tianyi, et al.
Pubblicazione: (2024)
di: Liu, Tianyi, et al.
Pubblicazione: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
di: Liao, Pan, et al.
Pubblicazione: (2026)
di: Liao, Pan, et al.
Pubblicazione: (2026)
Enhancing Incomplete Multi-modal Brain Tumor Segmentation with Intra-modal Asymmetry and Inter-modal Dependency
di: Liu, Weide, et al.
Pubblicazione: (2024)
di: Liu, Weide, et al.
Pubblicazione: (2024)
Why does Knowledge Distillation Work? Rethink its Attention and Fidelity Mechanism
di: Guo, Chenqi, et al.
Pubblicazione: (2024)
di: Guo, Chenqi, et al.
Pubblicazione: (2024)
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
di: Dagan, Gautier, et al.
Pubblicazione: (2024)
di: Dagan, Gautier, et al.
Pubblicazione: (2024)
MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty
di: Brödermann, Tim, et al.
Pubblicazione: (2024)
di: Brödermann, Tim, et al.
Pubblicazione: (2024)
Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation
di: Tao, Yitian, et al.
Pubblicazione: (2024)
di: Tao, Yitian, et al.
Pubblicazione: (2024)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
di: Li, Yunheng, et al.
Pubblicazione: (2024)
di: Li, Yunheng, et al.
Pubblicazione: (2024)
Multi-modal Attribute Prompting for Vision-Language Models
di: Liu, Xin, et al.
Pubblicazione: (2024)
di: Liu, Xin, et al.
Pubblicazione: (2024)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
di: Wu, Ruijia, et al.
Pubblicazione: (2025)
di: Wu, Ruijia, et al.
Pubblicazione: (2025)
Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment
di: Li, Jiaqing, et al.
Pubblicazione: (2026)
di: Li, Jiaqing, et al.
Pubblicazione: (2026)
From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts
di: Li, Weiran, et al.
Pubblicazione: (2025)
di: Li, Weiran, et al.
Pubblicazione: (2025)
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
di: Park, Sungjune, et al.
Pubblicazione: (2025)
di: Park, Sungjune, et al.
Pubblicazione: (2025)
PISE: Physics-Anchored Semantically-Enhanced Deep Computational Ghost Imaging for Robust Low-Bandwidth Machine Perception
di: Wu, Tong
Pubblicazione: (2026)
di: Wu, Tong
Pubblicazione: (2026)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
di: Cheng, Zihui, et al.
Pubblicazione: (2024)
di: Cheng, Zihui, et al.
Pubblicazione: (2024)
UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models
di: Liao, Zehui, et al.
Pubblicazione: (2025)
di: Liao, Zehui, et al.
Pubblicazione: (2025)
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
di: Hosseinzadeh, Mehdi, et al.
Pubblicazione: (2024)
di: Hosseinzadeh, Mehdi, et al.
Pubblicazione: (2024)
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction
di: Wei, Yujie, et al.
Pubblicazione: (2026)
di: Wei, Yujie, et al.
Pubblicazione: (2026)
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
di: Ma, Qihang, et al.
Pubblicazione: (2025)
di: Ma, Qihang, et al.
Pubblicazione: (2025)
Channel-adaptive Cross-modal Generative Semantic Communication for Point Cloud Transmission
di: Yang, Wanting, et al.
Pubblicazione: (2025)
di: Yang, Wanting, et al.
Pubblicazione: (2025)
MMLF: Multi-modal Multi-class Late Fusion for Object Detection with Uncertainty Estimation
di: Yang, Qihang, et al.
Pubblicazione: (2024)
di: Yang, Qihang, et al.
Pubblicazione: (2024)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
Deep Optimal Transport for Domain Adaptation on SPD Manifolds
di: Ju, Ce, et al.
Pubblicazione: (2022)
di: Ju, Ce, et al.
Pubblicazione: (2022)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
di: Liu, Haowei, et al.
Pubblicazione: (2024)
di: Liu, Haowei, et al.
Pubblicazione: (2024)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
di: Liang, Qiao, et al.
Pubblicazione: (2025)
di: Liang, Qiao, et al.
Pubblicazione: (2025)
Multi-Grained Cross-modal Alignment for Learning Open-vocabulary Semantic Segmentation from Text Supervision
di: Liu, Yajie, et al.
Pubblicazione: (2024)
di: Liu, Yajie, et al.
Pubblicazione: (2024)
Spatial-VLN: Zero-Shot Vision-and-Language Navigation With Explicit Spatial Perception and Exploration
di: Yue, Lu, et al.
Pubblicazione: (2026)
di: Yue, Lu, et al.
Pubblicazione: (2026)
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
di: Liu, Qinying, et al.
Pubblicazione: (2023)
di: Liu, Qinying, et al.
Pubblicazione: (2023)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
di: Jeon, Byungwoo, et al.
Pubblicazione: (2026)
di: Jeon, Byungwoo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Explaining Multi-modal Large Language Models by Analyzing their Vision Perception
di: Giulivi, Loris, et al.
Pubblicazione: (2024) -
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
di: Zhang, Ming, et al.
Pubblicazione: (2024) -
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
di: Miao, Yongzhu, et al.
Pubblicazione: (2023) -
APEX: Learning Adaptive Priorities for Multi-Objective Alignment in Vision-Language Generation
di: Chen, Dongliang, et al.
Pubblicazione: (2026) -
Hierarchical Vision-Language Interaction for Facial Action Unit Detection
di: Li, Yong, et al.
Pubblicazione: (2026)