SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Chunhachatrachai, Pawat, Faure, Gueter Josmy, Su, Hung-Ting, Hsu, Winston H. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
by: Chen, Posheng, et al.
Published: (2026)
by: Chen, Posheng, et al.
Published: (2026)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
by: Faure, Gueter Josmy, et al.
Published: (2026)
by: Faure, Gueter Josmy, et al.
Published: (2026)
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
by: Faure, Gueter Josmy, et al.
Published: (2024)
by: Faure, Gueter Josmy, et al.
Published: (2024)
MovieCORE: COgnitive REasoning in Movies
by: Faure, Gueter Josmy, et al.
Published: (2025)
by: Faure, Gueter Josmy, et al.
Published: (2025)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
by: Huang, Wei-Jhe, et al.
Published: (2024)
by: Huang, Wei-Jhe, et al.
Published: (2024)
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
by: Taguchi, Shun, et al.
Published: (2025)
by: Taguchi, Shun, et al.
Published: (2025)
Evolving, Not Training: Zero-Shot Reasoning Segmentation via Evolutionary Prompting
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
by: Liu, Ting, et al.
Published: (2025)
by: Liu, Ting, et al.
Published: (2025)
Prompt-Based Continual Compositional Zero-Shot Learning
by: Maryam, Sauda, et al.
Published: (2025)
by: Maryam, Sauda, et al.
Published: (2025)
Zero-Shot Industrial Anomaly Segmentation with Image-Aware Prompt Generation
by: Park, SoYoung, et al.
Published: (2025)
by: Park, SoYoung, et al.
Published: (2025)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
by: Zhang, Pu, et al.
Published: (2025)
by: Zhang, Pu, et al.
Published: (2025)
SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection
by: Fu, Chenhao, et al.
Published: (2026)
by: Fu, Chenhao, et al.
Published: (2026)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
by: Chen, Pei-An, et al.
Published: (2026)
by: Chen, Pei-An, et al.
Published: (2026)
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
GRAZE: Grounded Refinement and Motion-Aware Zero-Shot Event Localization
by: Zaidi, Syed Ahsan Masud, et al.
Published: (2026)
by: Zaidi, Syed Ahsan Masud, et al.
Published: (2026)
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
by: Byun, Ji Young, et al.
Published: (2025)
by: Byun, Ji Young, et al.
Published: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Transductive Zero-Shot and Few-Shot CLIP
by: Martin, Ségolène, et al.
Published: (2024)
by: Martin, Ségolène, et al.
Published: (2024)
HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images
by: Thota, Kundan, et al.
Published: (2026)
by: Thota, Kundan, et al.
Published: (2026)
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
by: Yang, Ziteng, et al.
Published: (2025)
by: Yang, Ziteng, et al.
Published: (2025)
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
by: Cai, Huanqia, et al.
Published: (2025)
by: Cai, Huanqia, et al.
Published: (2025)
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
by: Liao, Zhenyi, et al.
Published: (2025)
by: Liao, Zhenyi, et al.
Published: (2025)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025)
by: Zantout, Nader, et al.
Published: (2025)
MAGIC: Few-Shot Mask-Guided Anomaly Inpainting with Prompt Perturbation, Spatially Adaptive Guidance, and Context Awareness
by: Choi, JaeHyuck, et al.
Published: (2025)
by: Choi, JaeHyuck, et al.
Published: (2025)
GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow
by: Zeng, Ziyue, et al.
Published: (2026)
by: Zeng, Ziyue, et al.
Published: (2026)
DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial Editability
by: Hu, Xirui, et al.
Published: (2025)
by: Hu, Xirui, et al.
Published: (2025)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
by: Xu, Sirui, et al.
Published: (2024)
by: Xu, Sirui, et al.
Published: (2024)
Binary Verification for Zero-Shot Vision
by: Hu, Rongbin, et al.
Published: (2025)
by: Hu, Rongbin, et al.
Published: (2025)
VersusDebias: Universal Zero-Shot Debiasing for Text-to-Image Models via SLM-Based Prompt Engineering and Generative Adversary
by: Luo, Hanjun, et al.
Published: (2024)
by: Luo, Hanjun, et al.
Published: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
by: Jiang, Ruixiang, et al.
Published: (2025)
by: Jiang, Ruixiang, et al.
Published: (2025)
RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
by: Li, Ke, et al.
Published: (2025)
by: Li, Ke, et al.
Published: (2025)
Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
by: Wu, Aodi, et al.
Published: (2025)
by: Wu, Aodi, et al.
Published: (2025)
Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting
by: Frisoni, Giacomo, et al.
Published: (2026)
by: Frisoni, Giacomo, et al.
Published: (2026)
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
by: Han, Haonan, et al.
Published: (2026)
by: Han, Haonan, et al.
Published: (2026)
Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning
by: Lang, Jian, et al.
Published: (2025)
by: Lang, Jian, et al.
Published: (2025)
Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery
by: Mansour, Islam, et al.
Published: (2026)
by: Mansour, Islam, et al.
Published: (2026)
Coherent Zero-Shot Visual Instruction Generation
by: Phung, Quynh, et al.
Published: (2024)
by: Phung, Quynh, et al.
Published: (2024)
Similar Items
-
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
by: Chen, Posheng, et al.
Published: (2026) -
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
by: Faure, Gueter Josmy, et al.
Published: (2026) -
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
by: Faure, Gueter Josmy, et al.
Published: (2024) -
MovieCORE: COgnitive REasoning in Movies
by: Faure, Gueter Josmy, et al.
Published: (2025) -
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
by: Huang, Wei-Jhe, et al.
Published: (2024)