From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Oi, Masanari, Maeda, Koki, Koike, Ryuto, Oba, Daisuke, Inoue, Nakamasa, Okazaki, Naoaki |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024)
by: Oi, Masanari, et al.
Published: (2024)
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
by: Ichinose, Tatsuya, et al.
Published: (2026)
by: Ichinose, Tatsuya, et al.
Published: (2026)
Autoregressive Direct Preference Optimization
by: Oi, Masanari, et al.
Published: (2026)
by: Oi, Masanari, et al.
Published: (2026)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026)
by: Maeda, Koki, et al.
Published: (2026)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
by: Inoue, Nakamasa, et al.
Published: (2025)
by: Inoue, Nakamasa, et al.
Published: (2025)
Drifting Objectives for Refining Discrete Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Diffusion-State Policy Optimization for Masked Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
LLM Output Detectability and Task Performance Can be Jointly Optimized
by: Saito, Koshiro, et al.
Published: (2026)
by: Saito, Koshiro, et al.
Published: (2026)
JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
by: Sugiura, Issa, et al.
Published: (2026)
by: Sugiura, Issa, et al.
Published: (2026)
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
by: Koike, Ryuto, et al.
Published: (2025)
by: Koike, Ryuto, et al.
Published: (2025)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
by: Miyamoto, Sora, et al.
Published: (2026)
by: Miyamoto, Sora, et al.
Published: (2026)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
by: Sugiura, Issa, et al.
Published: (2026)
by: Sugiura, Issa, et al.
Published: (2026)
Machine Text Detectors are Membership Inference Attacks
by: Koike, Ryuto, et al.
Published: (2025)
by: Koike, Ryuto, et al.
Published: (2025)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model
by: Sasagawa, Keito, et al.
Published: (2024)
by: Sasagawa, Keito, et al.
Published: (2024)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
by: Ukai, Mahiro, et al.
Published: (2025)
by: Ukai, Mahiro, et al.
Published: (2025)
Multi-Point Positional Insertion Tuning for Small Object Detection
by: Goto, Kanoko, et al.
Published: (2024)
by: Goto, Kanoko, et al.
Published: (2024)
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
by: Mou, Tingshu, et al.
Published: (2026)
by: Mou, Tingshu, et al.
Published: (2026)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
by: Shen, Ruoyue, et al.
Published: (2024)
by: Shen, Ruoyue, et al.
Published: (2024)
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
by: Feng, Jie, et al.
Published: (2025)
by: Feng, Jie, et al.
Published: (2025)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)
by: Chen, Jiabin, et al.
Published: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
Multi-level and Multi-modal Action Anticipation
by: Kim, Seulgi, et al.
Published: (2025)
by: Kim, Seulgi, et al.
Published: (2025)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
by: Nishimura, Yuto, et al.
Published: (2024)
by: Nishimura, Yuto, et al.
Published: (2024)
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
by: Liu, Chonghan, et al.
Published: (2025)
by: Liu, Chonghan, et al.
Published: (2025)
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
by: Yasuki, Shunsuke, et al.
Published: (2025)
by: Yasuki, Shunsuke, et al.
Published: (2025)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
by: Xu, Weiye, et al.
Published: (2025)
by: Xu, Weiye, et al.
Published: (2025)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
by: Jiang, Dongsheng, et al.
Published: (2023)
by: Jiang, Dongsheng, et al.
Published: (2023)
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
by: Sugiura, Issa, et al.
Published: (2026)
by: Sugiura, Issa, et al.
Published: (2026)
Similar Items
-
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024) -
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024) -
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
by: Ichinose, Tatsuya, et al.
Published: (2026) -
Autoregressive Direct Preference Optimization
by: Oi, Masanari, et al.
Published: (2026) -
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026)