Visuospatial Perspective Taking in Multimodal Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Prunty, Jonathan, Zhang, Seraphina, Quinn, Patrick, Lian, Jianxun, Xie, Xing, Cheke, Lucy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
by: Hu, Zhengyu, et al.
Published: (2025)
by: Hu, Zhengyu, et al.
Published: (2025)
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
by: Yong, Xixian, et al.
Published: (2025)
by: Yong, Xixian, et al.
Published: (2025)
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
by: Burden, John, et al.
Published: (2025)
by: Burden, John, et al.
Published: (2025)
GraphInstruct: Empowering Large Language Models with Graph Understanding and Reasoning Capability
by: Luo, Zihan, et al.
Published: (2024)
by: Luo, Zihan, et al.
Published: (2024)
To Think or Not To Think, That is The Question for Large Reasoning Models in Theory of Mind Tasks
by: Gong, Nanxu, et al.
Published: (2026)
by: Gong, Nanxu, et al.
Published: (2026)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment
by: Mecattaf, Matteo G., et al.
Published: (2024)
by: Mecattaf, Matteo G., et al.
Published: (2024)
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
by: Chen, Hongyu, et al.
Published: (2025)
by: Chen, Hongyu, et al.
Published: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Visuospatial Cognitive Assistant
by: Feng, Qi
Published: (2025)
by: Feng, Qi
Published: (2025)
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
by: Fong, Seraphina, et al.
Published: (2025)
by: Fong, Seraphina, et al.
Published: (2025)
MultiContrievers: Analysis of Dense Retrieval Representations
by: Goldfarb-Tarrant, Seraphina, et al.
Published: (2024)
by: Goldfarb-Tarrant, Seraphina, et al.
Published: (2024)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025)
by: Yan, Shaotian, et al.
Published: (2025)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
by: Hildebrandt, Fabian, et al.
Published: (2025)
by: Hildebrandt, Fabian, et al.
Published: (2025)
The Good, The Bad, and Why: Unveiling Emotions in Generative AI
by: Li, Cheng, et al.
Published: (2023)
by: Li, Cheng, et al.
Published: (2023)
Multimodal Forecasting of Sparse Intraoperative Hypotension Events Powered by Language Model
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
UniToMBench: Integrating Perspective-Taking to Improve Theory of Mind in LLMs
by: Thiyagarajan, Prameshwar, et al.
Published: (2025)
by: Thiyagarajan, Prameshwar, et al.
Published: (2025)
UniMEL: A Unified Framework for Multimodal Entity Linking with Large Language Models
by: Qi, Liu, et al.
Published: (2024)
by: Qi, Liu, et al.
Published: (2024)
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
by: Jia, Boyu, et al.
Published: (2025)
by: Jia, Boyu, et al.
Published: (2025)
CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions
by: Rao, Jun, et al.
Published: (2024)
by: Rao, Jun, et al.
Published: (2024)
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
by: Hu, He, et al.
Published: (2025)
by: Hu, He, et al.
Published: (2025)
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
by: Ye, Yuxuan, et al.
Published: (2026)
by: Ye, Yuxuan, et al.
Published: (2026)
Can Language Models Take A Hint? Prompting for Controllable Contextualized Commonsense Inference
by: Colon-Hernandez, Pedro, et al.
Published: (2024)
by: Colon-Hernandez, Pedro, et al.
Published: (2024)
Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology
by: Liu, Dou, et al.
Published: (2025)
by: Liu, Dou, et al.
Published: (2025)
InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning
by: Xie, Congkai, et al.
Published: (2025)
by: Xie, Congkai, et al.
Published: (2025)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
by: Yuan, Youliang, et al.
Published: (2025)
by: Yuan, Youliang, et al.
Published: (2025)
Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
Population-Aligned Persona Generation for LLM-based Social Simulation
by: Hu, Zhengyu, et al.
Published: (2025)
by: Hu, Zhengyu, et al.
Published: (2025)
Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts
by: Feng, Qi
Published: (2025)
by: Feng, Qi
Published: (2025)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
by: Yi, Jingwei, et al.
Published: (2023)
by: Yi, Jingwei, et al.
Published: (2023)
EmoStage: A Framework for Accurate Empathetic Response Generation via Perspective-Taking and Phase Recognition
by: Qi, Zhiyang, et al.
Published: (2025)
by: Qi, Zhiyang, et al.
Published: (2025)
Decision-Level Ordinal Modeling for Multimodal Essay Scoring with Large Language Models
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models
by: Zhang, Che, et al.
Published: (2024)
by: Zhang, Che, et al.
Published: (2024)
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
by: Zhang, Yazhou, et al.
Published: (2024)
by: Zhang, Yazhou, et al.
Published: (2024)
Similar Items
-
Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
by: Hu, Zhengyu, et al.
Published: (2025) -
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
by: Yong, Xixian, et al.
Published: (2025) -
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
by: Burden, John, et al.
Published: (2025) -
GraphInstruct: Empowering Large Language Models with Graph Understanding and Reasoning Capability
by: Luo, Zihan, et al.
Published: (2024) -
To Think or Not To Think, That is The Question for Large Reasoning Models in Theory of Mind Tasks
by: Gong, Nanxu, et al.
Published: (2026)