SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Peiwen, Lang, Shiqiang, Wu, Dongming, Ding, Yi, Feng, Kaituo, Liu, Huadai, Ye, Zhen, Liu, Rui, Liu, Yun-Hui, Wang, Jianan, Yue, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Web to Pixels: Bringing Agentic Search into Visual Perception
by: Yang, Bokang, et al.
Published: (2026)
by: Yang, Bokang, et al.
Published: (2026)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
by: Sun, Peiwen, et al.
Published: (2026)
by: Sun, Peiwen, et al.
Published: (2026)
Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
by: Sun, Peiwen, et al.
Published: (2024)
by: Sun, Peiwen, et al.
Published: (2024)
OneThinker: All-in-one Reasoning Model for Image and Video
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
by: Wang, Yikun, et al.
Published: (2025)
by: Wang, Yikun, et al.
Published: (2025)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
by: Wu, Junfei, et al.
Published: (2025)
by: Wu, Junfei, et al.
Published: (2025)
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
by: Fan, Kaixuan, et al.
Published: (2025)
by: Fan, Kaixuan, et al.
Published: (2025)
LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
by: Xiao, Yijia, et al.
Published: (2024)
by: Xiao, Yijia, et al.
Published: (2024)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
by: Lu, Pan, et al.
Published: (2023)
by: Lu, Pan, et al.
Published: (2023)
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
by: Li, Lingxiao, et al.
Published: (2025)
by: Li, Lingxiao, et al.
Published: (2025)
AntCritic: Argument Mining for Free-Form and Visually-Rich Financial Comments
by: Liu, Huadai, et al.
Published: (2022)
by: Liu, Huadai, et al.
Published: (2022)
SR-Mamba: Effective Surgical Phase Recognition with State Space Model
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning
by: Chen, Li-Heng, et al.
Published: (2026)
by: Chen, Li-Heng, et al.
Published: (2026)
VistaDream: Sampling multiview consistent images for single-view scene reconstruction
by: Wang, Haiping, et al.
Published: (2024)
by: Wang, Haiping, et al.
Published: (2024)
OmniAudio: Generating Spatial Audio from 360-Degree Video
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
In Situ Nanoscale Probing of Lithium‐Aluminum Alloying / De‐Alloying Kinetics and Mechanical Failure in All‐Solid‐State Batteries
by: Rui‐Zhi Liu, et al.
Published: (2025)
by: Rui‐Zhi Liu, et al.
Published: (2025)
In Situ Nanoscale Probing of Lithium‐Aluminum Alloying / De‐Alloying Kinetics and Mechanical Failure in All‐Solid‐State Batteries
by: Rui‐Zhi Liu, et al.
Published: (2025)
by: Rui‐Zhi Liu, et al.
Published: (2025)
Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
by: He, Shuwei, et al.
Published: (2024)
by: He, Shuwei, et al.
Published: (2024)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
by: Zhu, Jiashun, et al.
Published: (2026)
by: Zhu, Jiashun, et al.
Published: (2026)
Structure Over Scale: Learning Visual Reasoning from Pedagogical Video
by: Galoaa, Bishoy, et al.
Published: (2026)
by: Galoaa, Bishoy, et al.
Published: (2026)
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
by: Yang, Zheyuan, et al.
Published: (2026)
by: Yang, Zheyuan, et al.
Published: (2026)
Exploring Reasoning Reward Model for Agents
by: Fan, Kaixuan, et al.
Published: (2026)
by: Fan, Kaixuan, et al.
Published: (2026)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
by: Su, Zhaochen, et al.
Published: (2026)
by: Su, Zhaochen, et al.
Published: (2026)
Devices, Functions, and Applications of Artificial Neuromorphic Visual Systems
by: Jiaxin Liu, et al.
Published: (2025)
by: Jiaxin Liu, et al.
Published: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
by: Liu, Huadai, et al.
Published: (2023)
by: Liu, Huadai, et al.
Published: (2023)
Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents
by: Zhang, Kaituo, et al.
Published: (2026)
by: Zhang, Kaituo, et al.
Published: (2026)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
Explore the Limits of Omni-modal Pretraining at Scale
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Advanced Distribution Theory for Significance in Scale Space
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Spatially Parallel All-optical Neural Networks
by: Qin, Jianwei, et al.
Published: (2025)
by: Qin, Jianwei, et al.
Published: (2025)
Hydrogen Bond Competition Optimizing Aqueous Zn Ion Solvation and (002) Interfacial Deposition with Ultralong Stability
by: Zhe Xiao, et al.
Published: (2025)
by: Zhe Xiao, et al.
Published: (2025)
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning
by: Jiang, Haoyi, et al.
Published: (2026)
by: Jiang, Haoyi, et al.
Published: (2026)
Experimental Demonstration of Twin-Field Quantum Digital Signatures over 504 km
by: Zhang, Chun-Hui, et al.
Published: (2026)
by: Zhang, Chun-Hui, et al.
Published: (2026)
Similar Items
-
From Web to Pixels: Bringing Agentic Search into Visual Perception
by: Yang, Bokang, et al.
Published: (2026) -
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
by: Sun, Peiwen, et al.
Published: (2026) -
Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
by: Sun, Peiwen, et al.
Published: (2024) -
OneThinker: All-in-one Reasoning Model for Image and Video
by: Feng, Kaituo, et al.
Published: (2025) -
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
by: Wang, Yikun, et al.
Published: (2025)