Are MLMs Trapped in the Visual Room?
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yazhou, Zou, Chunwang, Liu, Qimeng, Rong, Lu, Yao, Ben, Lian, Zheng, Li, Qiuchi, Zhang, Peng, Qin, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
di: Zhang, Yazhou, et al.
Pubblicazione: (2025)
di: Zhang, Yazhou, et al.
Pubblicazione: (2025)
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
di: Zhang, Yuchen, et al.
Pubblicazione: (2025)
di: Zhang, Yuchen, et al.
Pubblicazione: (2025)
LXLv2: Enhanced LiDAR Excluded Lean 3D Object Detection with Fusion of 4D Radar and Camera
di: Xiong, Weiyi, et al.
Pubblicazione: (2025)
di: Xiong, Weiyi, et al.
Pubblicazione: (2025)
AirRoom: Objects Matter in Room Reidentification
di: Yao, Runmao, et al.
Pubblicazione: (2025)
di: Yao, Runmao, et al.
Pubblicazione: (2025)
MAUP: Training-free Multi-center Adaptive Uncertainty-aware Prompting for Cross-domain Few-shot Medical Image Segmentation
di: Zhu, Yazhou, et al.
Pubblicazione: (2025)
di: Zhu, Yazhou, et al.
Pubblicazione: (2025)
ESR-NeRF: Emissive Source Reconstruction Using LDR Multi-view Images
di: Jeong, Jinseo, et al.
Pubblicazione: (2024)
di: Jeong, Jinseo, et al.
Pubblicazione: (2024)
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
di: Yao, Ruilin, et al.
Pubblicazione: (2026)
di: Yao, Ruilin, et al.
Pubblicazione: (2026)
Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition
di: Liu, Ran, et al.
Pubblicazione: (2025)
di: Liu, Ran, et al.
Pubblicazione: (2025)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
di: Chang, Boyu, et al.
Pubblicazione: (2026)
di: Chang, Boyu, et al.
Pubblicazione: (2026)
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
Dynamic in Static: Hybrid Visual Correspondence for Self-Supervised Video Object Segmentation
di: Pei, Gensheng, et al.
Pubblicazione: (2024)
di: Pei, Gensheng, et al.
Pubblicazione: (2024)
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
di: Yao, Jianhang, et al.
Pubblicazione: (2025)
di: Yao, Jianhang, et al.
Pubblicazione: (2025)
Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps
di: Zhang, Jianwei, et al.
Pubblicazione: (2026)
di: Zhang, Jianwei, et al.
Pubblicazione: (2026)
PolyRoom: Room-aware Transformer for Floorplan Reconstruction
di: Liu, Yuzhou, et al.
Pubblicazione: (2024)
di: Liu, Yuzhou, et al.
Pubblicazione: (2024)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
di: Zhang, Wang, et al.
Pubblicazione: (2024)
di: Zhang, Wang, et al.
Pubblicazione: (2024)
CAR: Controllable Autoregressive Modeling for Visual Generation
di: Yao, Ziyu, et al.
Pubblicazione: (2024)
di: Yao, Ziyu, et al.
Pubblicazione: (2024)
FAMNet: Frequency-aware Matching Network for Cross-domain Few-shot Medical Image Segmentation
di: Bo, Yuntian, et al.
Pubblicazione: (2024)
di: Bo, Yuntian, et al.
Pubblicazione: (2024)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
di: Zhang, Qizhe, et al.
Pubblicazione: (2023)
di: Zhang, Qizhe, et al.
Pubblicazione: (2023)
Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?
di: Yao, Ben, et al.
Pubblicazione: (2024)
di: Yao, Ben, et al.
Pubblicazione: (2024)
VisionTrap: Unanswerable Questions On Visual Data
di: Saadat, Asir, et al.
Pubblicazione: (2025)
di: Saadat, Asir, et al.
Pubblicazione: (2025)
Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?
di: Han, Cheng, et al.
Pubblicazione: (2024)
di: Han, Cheng, et al.
Pubblicazione: (2024)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
di: Zhang, Ben, et al.
Pubblicazione: (2025)
di: Zhang, Ben, et al.
Pubblicazione: (2025)
SEP: Self-Enhanced Prompt Tuning for Visual-Language Model
di: Yao, Hantao, et al.
Pubblicazione: (2024)
di: Yao, Hantao, et al.
Pubblicazione: (2024)
TDANet: Target-Directed Attention Network For Object-Goal Visual Navigation With Zero-Shot Ability
di: Lian, Shiwei, et al.
Pubblicazione: (2024)
di: Lian, Shiwei, et al.
Pubblicazione: (2024)
Combating Noisy Labels through Fostering Self- and Neighbor-Consistency
di: Sun, Zeren, et al.
Pubblicazione: (2026)
di: Sun, Zeren, et al.
Pubblicazione: (2026)
Mind the Discriminability Trap in Source-Free Cross-domain Few-shot Learning
di: Zhang, Zhenyu, et al.
Pubblicazione: (2026)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2026)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
di: Li, Haokun, et al.
Pubblicazione: (2025)
di: Li, Haokun, et al.
Pubblicazione: (2025)
Evaluating Time Awareness and Cross-modal Active Perception of Large Models via 4D Escape Room Task
di: Dong, Yurui, et al.
Pubblicazione: (2026)
di: Dong, Yurui, et al.
Pubblicazione: (2026)
COMOGen: A Controllable Text-to-3D Multi-object Generation Framework
di: Sun, Shaorong, et al.
Pubblicazione: (2024)
di: Sun, Shaorong, et al.
Pubblicazione: (2024)
Spatial-ORMLLM: Improve Spatial Relation Understanding in the Operating Room with Multimodal Large Language Model
di: He, Peiqi, et al.
Pubblicazione: (2025)
di: He, Peiqi, et al.
Pubblicazione: (2025)
FTMoMamba: Motion Generation with Frequency and Text State Space Models
di: Li, Chengjian, et al.
Pubblicazione: (2024)
di: Li, Chengjian, et al.
Pubblicazione: (2024)
Twofold Debiasing Enhances Fine-Grained Learning with Coarse Labels
di: Zhao, Xin-yang, et al.
Pubblicazione: (2025)
di: Zhao, Xin-yang, et al.
Pubblicazione: (2025)
CauSight: Learning to Supersense for Visual Causal Discovery
di: Zhang, Yize, et al.
Pubblicazione: (2025)
di: Zhang, Yize, et al.
Pubblicazione: (2025)
Training A Small Emotional Vision Language Model for Visual Art Comprehension
di: Zhang, Jing, et al.
Pubblicazione: (2024)
di: Zhang, Jing, et al.
Pubblicazione: (2024)
Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval
di: Zhang, Jing, et al.
Pubblicazione: (2026)
di: Zhang, Jing, et al.
Pubblicazione: (2026)
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
di: Zhang, Yazhou, et al.
Pubblicazione: (2024)
di: Zhang, Yazhou, et al.
Pubblicazione: (2024)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
di: Zou, Xin, et al.
Pubblicazione: (2025)
di: Zou, Xin, et al.
Pubblicazione: (2025)
Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting
di: Bo, Yuntian, et al.
Pubblicazione: (2026)
di: Bo, Yuntian, et al.
Pubblicazione: (2026)
VIALM: A Survey and Benchmark of Visually Impaired Assistance with Large Models
di: Zhao, Yi, et al.
Pubblicazione: (2024)
di: Zhao, Yi, et al.
Pubblicazione: (2024)
Prim2Room: Layout-Controllable Room Mesh Generation from Primitives
di: Feng, Chengzeng, et al.
Pubblicazione: (2024)
di: Feng, Chengzeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
di: Zhang, Yazhou, et al.
Pubblicazione: (2025) -
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
di: Zhang, Yuchen, et al.
Pubblicazione: (2025) -
LXLv2: Enhanced LiDAR Excluded Lean 3D Object Detection with Fusion of 4D Radar and Camera
di: Xiong, Weiyi, et al.
Pubblicazione: (2025) -
AirRoom: Objects Matter in Room Reidentification
di: Yao, Runmao, et al.
Pubblicazione: (2025) -
MAUP: Training-free Multi-center Adaptive Uncertainty-aware Prompting for Cross-domain Few-shot Medical Image Segmentation
di: Zhu, Yazhou, et al.
Pubblicazione: (2025)