Salvato in:
| Autori principali: | Lou, Haoran, Liu, Ziyan, Fan, Chunxiao, Wu, Yuexin, Ming, Yue, Wu, Hao, Zuo, Kai, Chen, Yibo, Tang, Xu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.13710 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
di: Lou, Haoran, et al.
Pubblicazione: (2025)
di: Lou, Haoran, et al.
Pubblicazione: (2025)
MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
di: Liu, Ziyan, et al.
Pubblicazione: (2025)
di: Liu, Ziyan, et al.
Pubblicazione: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
di: Tong, Jintao, et al.
Pubblicazione: (2025)
di: Tong, Jintao, et al.
Pubblicazione: (2025)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
di: Wang, Zitian, et al.
Pubblicazione: (2025)
di: Wang, Zitian, et al.
Pubblicazione: (2025)
Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality
di: Qian, Kai, et al.
Pubblicazione: (2026)
di: Qian, Kai, et al.
Pubblicazione: (2026)
Heterogeneous-Modal Unsupervised Domain Adaptation via Latent Space Bridging
di: Yang, Jiawen, et al.
Pubblicazione: (2025)
di: Yang, Jiawen, et al.
Pubblicazione: (2025)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
di: Zhou, Hefeng, et al.
Pubblicazione: (2026)
di: Zhou, Hefeng, et al.
Pubblicazione: (2026)
Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memory
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
ROG: Retrieval-Augmented LLM Reasoning for Complex First-Order Queries over Knowledge Graphs
di: Zhang, Ziyan, et al.
Pubblicazione: (2026)
di: Zhang, Ziyan, et al.
Pubblicazione: (2026)
Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges
di: Nguyen-Le, Hong-Hanh, et al.
Pubblicazione: (2026)
di: Nguyen-Le, Hong-Hanh, et al.
Pubblicazione: (2026)
Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain
di: Chao, Lianying, et al.
Pubblicazione: (2026)
di: Chao, Lianying, et al.
Pubblicazione: (2026)
Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Heads
di: Patel, Shaswat, et al.
Pubblicazione: (2026)
di: Patel, Shaswat, et al.
Pubblicazione: (2026)
Towards Convexity in Anomaly Detection: A New Formulation of SSLM with Unique Optimal Solutions
di: Liu, Hongying, et al.
Pubblicazione: (2024)
di: Liu, Hongying, et al.
Pubblicazione: (2024)
Growing Visual Generative Capacity for Pre-Trained MLLMs
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
Unveiling and Bridging the Functional Perception Gap in MLLMs: Atomic Visual Alignment and Hierarchical Evaluation via PET-Bench
di: Ye, Zanting, et al.
Pubblicazione: (2026)
di: Ye, Zanting, et al.
Pubblicazione: (2026)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
di: Zhang, Xu, et al.
Pubblicazione: (2025)
di: Zhang, Xu, et al.
Pubblicazione: (2025)
From Query to Explanation: Uni-RAG for Multi-Modal Retrieval-Augmented Learning in STEM
di: Wu, Xinyi, et al.
Pubblicazione: (2025)
di: Wu, Xinyi, et al.
Pubblicazione: (2025)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
di: Huang, Jincai, et al.
Pubblicazione: (2026)
di: Huang, Jincai, et al.
Pubblicazione: (2026)
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression
di: Li, Yuankai, et al.
Pubblicazione: (2024)
di: Li, Yuankai, et al.
Pubblicazione: (2024)
Bridging Search and Recommendation through Latent Cross Reasoning
di: Shi, Teng, et al.
Pubblicazione: (2025)
di: Shi, Teng, et al.
Pubblicazione: (2025)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Multi-Masked Querying Network for Robust Emotion Recognition from Incomplete Multi-Modal Physiological Signals
di: Xu, Geng-Xin, et al.
Pubblicazione: (2025)
di: Xu, Geng-Xin, et al.
Pubblicazione: (2025)
Graph Transfer Learning via Shared Latent Geometry: Theory and Applications
di: Wu, Tong, et al.
Pubblicazione: (2026)
di: Wu, Tong, et al.
Pubblicazione: (2026)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
MLLMs are Deeply Affected by Modality Bias
di: Zheng, Xu, et al.
Pubblicazione: (2025)
di: Zheng, Xu, et al.
Pubblicazione: (2025)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
di: Huang, Zhe, et al.
Pubblicazione: (2025)
di: Huang, Zhe, et al.
Pubblicazione: (2025)
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
di: Xue, Wangyu, et al.
Pubblicazione: (2024)
di: Xue, Wangyu, et al.
Pubblicazione: (2024)
VisualNeo: Bridging the Gap between Visual Query Interfaces and Graph Query Engines
di: Huang, Kai, et al.
Pubblicazione: (2026)
di: Huang, Kai, et al.
Pubblicazione: (2026)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
di: Zhou, Zhongzhu, et al.
Pubblicazione: (2026)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
di: Chen, Guizhen, et al.
Pubblicazione: (2025)
di: Chen, Guizhen, et al.
Pubblicazione: (2025)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
Bridging Queries and Tables through Entities in Table Retrieval
di: Li, Da, et al.
Pubblicazione: (2025)
di: Li, Da, et al.
Pubblicazione: (2025)
Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities
di: Fan, Qi, et al.
Pubblicazione: (2024)
di: Fan, Qi, et al.
Pubblicazione: (2024)
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
di: Zhang, Bob, et al.
Pubblicazione: (2025)
di: Zhang, Bob, et al.
Pubblicazione: (2025)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
di: E, Shaojun, et al.
Pubblicazione: (2025)
di: E, Shaojun, et al.
Pubblicazione: (2025)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
di: Wu, Hao, et al.
Pubblicazione: (2026)
di: Wu, Hao, et al.
Pubblicazione: (2026)
CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction
di: Ma, Hao, et al.
Pubblicazione: (2024)
di: Ma, Hao, et al.
Pubblicazione: (2024)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
di: Dai, Yifan, et al.
Pubblicazione: (2026)
di: Dai, Yifan, et al.
Pubblicazione: (2026)
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
di: Xu, Wenyuan, et al.
Pubblicazione: (2025)
di: Xu, Wenyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
di: Lou, Haoran, et al.
Pubblicazione: (2025) -
MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
di: Liu, Ziyan, et al.
Pubblicazione: (2025) -
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
di: Tong, Jintao, et al.
Pubblicazione: (2025) -
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
di: Wang, Zitian, et al.
Pubblicazione: (2025) -
Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality
di: Qian, Kai, et al.
Pubblicazione: (2026)