Mind over Space: Can Multimodal Large Language Models Mentally Navigate?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Qihui, Ruan, Shouwei, Yang, Xiao, Jiang, Hao, Huang, Yao, Zhao, Shiji, Fan, Hanwei, Su, Hang, Wei, Xingxing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack
von: Liu, Hanqing, et al.
Veröffentlicht: (2025)
von: Liu, Hanqing, et al.
Veröffentlicht: (2025)
World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models
von: Ruan, Shouwei, et al.
Veröffentlicht: (2026)
von: Ruan, Shouwei, et al.
Veröffentlicht: (2026)
Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models
von: Ruan, Shouwei, et al.
Veröffentlicht: (2024)
von: Ruan, Shouwei, et al.
Veröffentlicht: (2024)
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
Towards Transferable Targeted 3D Adversarial Attack in the Physical World
von: Huang, Yao, et al.
Veröffentlicht: (2023)
von: Huang, Yao, et al.
Veröffentlicht: (2023)
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?
von: Ruan, Shouwei, et al.
Veröffentlicht: (2024)
von: Ruan, Shouwei, et al.
Veröffentlicht: (2024)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
von: Huang, Yao, et al.
Veröffentlicht: (2025)
von: Huang, Yao, et al.
Veröffentlicht: (2025)
OODFace: Benchmarking Robustness of Face Recognition under Common Corruptions and Appearance Variations
von: Kang, Caixin, et al.
Veröffentlicht: (2024)
von: Kang, Caixin, et al.
Veröffentlicht: (2024)
Improving Safety Alignment via Balanced Direct Preference Optimization
von: Zhao, Shiji, et al.
Veröffentlicht: (2026)
von: Zhao, Shiji, et al.
Veröffentlicht: (2026)
DIFFender: Diffusion-Based Adversarial Defense against Patch Attacks
von: Kang, Caixin, et al.
Veröffentlicht: (2023)
von: Kang, Caixin, et al.
Veröffentlicht: (2023)
Real-world Adversarial Defense against Patch Attacks based on Diffusion Model
von: Wei, Xingxing, et al.
Veröffentlicht: (2024)
von: Wei, Xingxing, et al.
Veröffentlicht: (2024)
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
Exploring the Robustness of Decision-Level Through Adversarial Attacks on LLM-Based Embodied Models
von: Liu, Shuyuan, et al.
Veröffentlicht: (2024)
von: Liu, Shuyuan, et al.
Veröffentlicht: (2024)
NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation
von: Sun, Yitong, et al.
Veröffentlicht: (2025)
von: Sun, Yitong, et al.
Veröffentlicht: (2025)
Mitigating Accuracy-Robustness Trade-off via Balanced Multi-Teacher Adversarial Distillation
von: Zhao, Shiji, et al.
Veröffentlicht: (2023)
von: Zhao, Shiji, et al.
Veröffentlicht: (2023)
Revisiting the Trade-off between Accuracy and Robustness via Weight Distribution of Filters
von: Wei, Xingxing, et al.
Veröffentlicht: (2023)
von: Wei, Xingxing, et al.
Veröffentlicht: (2023)
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
von: Song, Dingjie, et al.
Veröffentlicht: (2026)
von: Song, Dingjie, et al.
Veröffentlicht: (2026)
GSVA: Generalized Segmentation via Multimodal Large Language Models
von: Xia, Zhuofan, et al.
Veröffentlicht: (2023)
von: Xia, Zhuofan, et al.
Veröffentlicht: (2023)
Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning
von: Huang, Zhuo, et al.
Veröffentlicht: (2023)
von: Huang, Zhuo, et al.
Veröffentlicht: (2023)
Improving Adversarial Robust Fairness via Anti-Bias Soft Label Distillation
von: Zhao, Shiji, et al.
Veröffentlicht: (2023)
von: Zhao, Shiji, et al.
Veröffentlicht: (2023)
ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
von: Huang, Zhenyang, et al.
Veröffentlicht: (2025)
von: Huang, Zhenyang, et al.
Veröffentlicht: (2025)
Regular rings over valuation rings
von: Lyu, Shiji
Veröffentlicht: (2026)
von: Lyu, Shiji
Veröffentlicht: (2026)
PrivacyMind: Large Language Models Can Be Contextual Privacy Protection Learners
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
LLM4Rec: Large Language Models for Multimodal Generative Recommendation with Causal Debiasing
von: Ma, Bo, et al.
Veröffentlicht: (2025)
von: Ma, Bo, et al.
Veröffentlicht: (2025)
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
RecMind: Large Language Model Powered Agent For Recommendation
von: Wang, Yancheng, et al.
Veröffentlicht: (2023)
von: Wang, Yancheng, et al.
Veröffentlicht: (2023)
A Comprehensive Survey of Continual Learning: Theory, Method and Application
von: Wang, Liyuan, et al.
Veröffentlicht: (2023)
von: Wang, Liyuan, et al.
Veröffentlicht: (2023)
Towards Class-wise Fair Adversarial Training via Anti-Bias Soft Label Distillation
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
Visual Adversarial Attacks and Defenses in the Physical World: A Survey
von: Wei, Xingxing, et al.
Veröffentlicht: (2022)
von: Wei, Xingxing, et al.
Veröffentlicht: (2022)
MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation
von: Zhang, Qihui, et al.
Veröffentlicht: (2025)
von: Zhang, Qihui, et al.
Veröffentlicht: (2025)
MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack
von: Liu, Hanqing, et al.
Veröffentlicht: (2025) -
World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models
von: Ruan, Shouwei, et al.
Veröffentlicht: (2026) -
Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models
von: Ruan, Shouwei, et al.
Veröffentlicht: (2024) -
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025) -
Towards Transferable Targeted 3D Adversarial Attack in the Physical World
von: Huang, Yao, et al.
Veröffentlicht: (2023)