See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Yuxi, Huang, Wei, Chen, Qirui, Hou, Lu, Qi, Xiaojuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
von: Long, Lin, et al.
Veröffentlicht: (2025)
von: Long, Lin, et al.
Veröffentlicht: (2025)
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
von: Xiao, Yicheng, et al.
Veröffentlicht: (2026)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2026)
Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models
von: Zhang, Baoheng, et al.
Veröffentlicht: (2026)
von: Zhang, Baoheng, et al.
Veröffentlicht: (2026)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
von: Deng, Wei, et al.
Veröffentlicht: (2026)
von: Deng, Wei, et al.
Veröffentlicht: (2026)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
von: Wang, Qirui, et al.
Veröffentlicht: (2026)
von: Wang, Qirui, et al.
Veröffentlicht: (2026)
Building Extraction from Remote Sensing Imagery under Hazy and Low-light Conditions: Benchmark and Baseline
von: Sang, Feifei, et al.
Veröffentlicht: (2026)
von: Sang, Feifei, et al.
Veröffentlicht: (2026)
SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements
von: Xie, Haiyang, et al.
Veröffentlicht: (2025)
von: Xie, Haiyang, et al.
Veröffentlicht: (2025)
ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models
von: Shen, Qirui, et al.
Veröffentlicht: (2026)
von: Shen, Qirui, et al.
Veröffentlicht: (2026)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
von: Hou, Wenjin, et al.
Veröffentlicht: (2026)
von: Hou, Wenjin, et al.
Veröffentlicht: (2026)
Multiple Object Tracking in Video SAR: A Benchmark and Tracking Baseline
von: Chen, Haoxiang, et al.
Veröffentlicht: (2025)
von: Chen, Haoxiang, et al.
Veröffentlicht: (2025)
VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
von: Liu, Pengyiang, et al.
Veröffentlicht: (2026)
von: Liu, Pengyiang, et al.
Veröffentlicht: (2026)
Event-based Tiny Object Detection: A Benchmark Dataset and Baseline
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
von: Wang, Chunwei, et al.
Veröffentlicht: (2024)
von: Wang, Chunwei, et al.
Veröffentlicht: (2024)
Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets
von: Wu, Peng, et al.
Veröffentlicht: (2026)
von: Wu, Peng, et al.
Veröffentlicht: (2026)
Multi-Modal Building Change Detection for Large-Scale Small Changes: Benchmark and Baseline
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
von: Huang, Xinmiao, et al.
Veröffentlicht: (2025)
von: Huang, Xinmiao, et al.
Veröffentlicht: (2025)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
Seeing the Unseen in Low-light Spike Streams
von: Hu, Liwen, et al.
Veröffentlicht: (2025)
von: Hu, Liwen, et al.
Veröffentlicht: (2025)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
von: Ma, Wenxin, et al.
Veröffentlicht: (2026)
von: Ma, Wenxin, et al.
Veröffentlicht: (2026)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
von: Zhang, Jialiang, et al.
Veröffentlicht: (2026)
von: Zhang, Jialiang, et al.
Veröffentlicht: (2026)
MeMix: Writing Less, Remembering More for Streaming 3D Reconstruction
von: Dong, Jiacheng, et al.
Veröffentlicht: (2026)
von: Dong, Jiacheng, et al.
Veröffentlicht: (2026)
Gait Recognition in the Wild: A Large-scale Benchmark and NAS-based Baseline
von: Guo, Xianda, et al.
Veröffentlicht: (2022)
von: Guo, Xianda, et al.
Veröffentlicht: (2022)
SPR-128K: A New Benchmark for Spatial Plausibility Reasoning with Multimodal Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
von: Wang, Qirui, et al.
Veröffentlicht: (2025)
von: Wang, Qirui, et al.
Veröffentlicht: (2025)
PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors
von: Yang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxin, et al.
Veröffentlicht: (2026)
OpenStereo: A Comprehensive Benchmark for Stereo Matching and Strong Baseline
von: Guo, Xianda, et al.
Veröffentlicht: (2023)
von: Guo, Xianda, et al.
Veröffentlicht: (2023)
Semantic Correspondence: Unified Benchmarking and a Strong Baseline
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2025)
You Only Speak Once to See
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
A Semi-supervised Nighttime Dehazing Baseline with Spatial-Frequency Aware and Realistic Brightness Constraint
von: Cong, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Cong, Xiaofeng, et al.
Veröffentlicht: (2024)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
von: Lu, Xudong, et al.
Veröffentlicht: (2026)
von: Lu, Xudong, et al.
Veröffentlicht: (2026)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
Real-World Remote Sensing Image Dehazing: Benchmark and Baseline
von: Zhu, Zeng-Hui, et al.
Veröffentlicht: (2025)
von: Zhu, Zeng-Hui, et al.
Veröffentlicht: (2025)
Learning Streaming Video Representation via Multitask Training
von: Yan, Yibin, et al.
Veröffentlicht: (2025)
von: Yan, Yibin, et al.
Veröffentlicht: (2025)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
von: Shi, Yudi, et al.
Veröffentlicht: (2024)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
von: Long, Lin, et al.
Veröffentlicht: (2025) -
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025) -
Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline
von: Wei, Tianqi, et al.
Veröffentlicht: (2024) -
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
von: Xiao, Yicheng, et al.
Veröffentlicht: (2026) -
Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models
von: Zhang, Baoheng, et al.
Veröffentlicht: (2026)