When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akasaka, Muku, Han, Soyeon Caren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
von: Xiang, Biao, et al.
Veröffentlicht: (2026)
von: Xiang, Biao, et al.
Veröffentlicht: (2026)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
Deep Learning based Visually Rich Document Content Understanding: A Survey
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
Graph Neural Networks for Text Classification: A Survey
von: Wang, Kunze, et al.
Veröffentlicht: (2023)
von: Wang, Kunze, et al.
Veröffentlicht: (2023)
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
von: Li, Yan, et al.
Veröffentlicht: (2024)
von: Li, Yan, et al.
Veröffentlicht: (2024)
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
von: Dai, Yue, et al.
Veröffentlicht: (2025)
von: Dai, Yue, et al.
Veröffentlicht: (2025)
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
MSG-Chart: Multimodal Scene Graph for ChartQA
von: Dai, Yue, et al.
Veröffentlicht: (2024)
von: Dai, Yue, et al.
Veröffentlicht: (2024)
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
von: Ng, Thye Shan, et al.
Veröffentlicht: (2024)
von: Ng, Thye Shan, et al.
Veröffentlicht: (2024)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
In-game Toxic Language Detection: Shared Task and Attention Residuals
von: Jia, Yuanzhe, et al.
Veröffentlicht: (2022)
von: Jia, Yuanzhe, et al.
Veröffentlicht: (2022)
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
von: Han, Soyeon Caren, et al.
Veröffentlicht: (2024)
von: Han, Soyeon Caren, et al.
Veröffentlicht: (2024)
LIMO: Less is More for Reasoning
von: Ye, Yixin, et al.
Veröffentlicht: (2025)
von: Ye, Yixin, et al.
Veröffentlicht: (2025)
A Survey of Large Language Models in Finance (FinLLMs)
von: Lee, Jean, et al.
Veröffentlicht: (2024)
von: Lee, Jean, et al.
Veröffentlicht: (2024)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
von: Park, Jiwon, et al.
Veröffentlicht: (2025)
von: Park, Jiwon, et al.
Veröffentlicht: (2025)
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
von: Li, Haoyi, et al.
Veröffentlicht: (2025)
von: Li, Haoyi, et al.
Veröffentlicht: (2025)
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
von: Cabral, Rina Carines, et al.
Veröffentlicht: (2024)
von: Cabral, Rina Carines, et al.
Veröffentlicht: (2024)
EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
von: Yuan, Angela Yifei, et al.
Veröffentlicht: (2025)
von: Yuan, Angela Yifei, et al.
Veröffentlicht: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU
von: Li, Yan, et al.
Veröffentlicht: (2024)
von: Li, Yan, et al.
Veröffentlicht: (2024)
Commonsense Reasoning in Arab Culture
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
von: Madani, Mohammad Reza Ghasemi, et al.
Veröffentlicht: (2026)
von: Madani, Mohammad Reza Ghasemi, et al.
Veröffentlicht: (2026)
Soft Prompt Tuning for Cross-Lingual Transfer: When Less is More
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
von: De Cristofaro, Domenico, et al.
Veröffentlicht: (2026)
von: De Cristofaro, Domenico, et al.
Veröffentlicht: (2026)
When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
von: Kapur, Rhea, et al.
Veröffentlicht: (2026)
von: Kapur, Rhea, et al.
Veröffentlicht: (2026)
Doing Good or Doing Right? Exploring the Weakness of Commonsense Causal Reasoning Models
von: Han, Mingyue, et al.
Veröffentlicht: (2021)
von: Han, Mingyue, et al.
Veröffentlicht: (2021)
When More is Less: Understanding Chain-of-Thought Length in LLMs
von: Wu, Yuyang, et al.
Veröffentlicht: (2025)
von: Wu, Yuyang, et al.
Veröffentlicht: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning
von: Sui, Yi, et al.
Veröffentlicht: (2026)
von: Sui, Yi, et al.
Veröffentlicht: (2026)
Acting Less is Reasoning More! Teaching Model to Act Efficiently
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
von: Wang, Hongru, et al.
Veröffentlicht: (2025)
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
von: Zhang, Qianchi, et al.
Veröffentlicht: (2025)
von: Zhang, Qianchi, et al.
Veröffentlicht: (2025)
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
von: Cui, Jin, et al.
Veröffentlicht: (2026)
von: Cui, Jin, et al.
Veröffentlicht: (2026)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
von: Zhao, Haoran, et al.
Veröffentlicht: (2026)
von: Zhao, Haoran, et al.
Veröffentlicht: (2026)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2024) -
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
von: Yang, Shuo, et al.
Veröffentlicht: (2025) -
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
von: Xiang, Biao, et al.
Veröffentlicht: (2026) -
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
von: Wang, Eileen, et al.
Veröffentlicht: (2024) -
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)