Visual Agentic AI for Spatial Reasoning with a Dynamic API
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Marsili, Damiano, Agrawal, Rohun, Yue, Yisong, Gkioxari, Georgia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
von: Marsili, Damiano, et al.
Veröffentlicht: (2025)
von: Marsili, Damiano, et al.
Veröffentlicht: (2025)
Same or Not? Enhancing Visual Perception in Vision-Language Models
von: Marsili, Damiano, et al.
Veröffentlicht: (2025)
von: Marsili, Damiano, et al.
Veröffentlicht: (2025)
Find Any Part in 3D
von: Ma, Ziqi, et al.
Veröffentlicht: (2024)
von: Ma, Ziqi, et al.
Veröffentlicht: (2024)
Feedforward 3D Editing via Text-Steerable Image-to-3D
von: Ma, Ziqi, et al.
Veröffentlicht: (2025)
von: Ma, Ziqi, et al.
Veröffentlicht: (2025)
Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
von: Demler, Ilona, et al.
Veröffentlicht: (2025)
von: Demler, Ilona, et al.
Veröffentlicht: (2025)
Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
von: Sahoo, Aadarsh, et al.
Veröffentlicht: (2026)
von: Sahoo, Aadarsh, et al.
Veröffentlicht: (2026)
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
von: Kang, Raphi, et al.
Veröffentlicht: (2026)
von: Kang, Raphi, et al.
Veröffentlicht: (2026)
Aligning Text, Images, and 3D Structure Token-by-Token
von: Sahoo, Aadarsh, et al.
Veröffentlicht: (2025)
von: Sahoo, Aadarsh, et al.
Veröffentlicht: (2025)
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
von: Ma, Ziqi, et al.
Veröffentlicht: (2026)
von: Ma, Ziqi, et al.
Veröffentlicht: (2026)
Is CLIP ideal? No. Can we fix it? Yes!
von: Kang, Raphi, et al.
Veröffentlicht: (2025)
von: Kang, Raphi, et al.
Veröffentlicht: (2025)
Reconstructing Hand-Held Objects in 3D from Images and Videos
von: Wu, Jane, et al.
Veröffentlicht: (2024)
von: Wu, Jane, et al.
Veröffentlicht: (2024)
MonoTher-Depth: Enhancing Thermal Depth Estimation via Confidence-Aware Distillation
von: Zuo, Xingxing, et al.
Veröffentlicht: (2025)
von: Zuo, Xingxing, et al.
Veröffentlicht: (2025)
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
von: Agrawal, Palaash, et al.
Veröffentlicht: (2023)
von: Agrawal, Palaash, et al.
Veröffentlicht: (2023)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
von: Yang, Qian, et al.
Veröffentlicht: (2026)
von: Yang, Qian, et al.
Veröffentlicht: (2026)
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2026)
von: Li, Yian, et al.
Veröffentlicht: (2026)
Caltech Aerial RGB-Thermal Dataset in the Wild
von: Lee, Connor, et al.
Veröffentlicht: (2024)
von: Lee, Connor, et al.
Veröffentlicht: (2024)
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
von: Sun, Peiwen, et al.
Veröffentlicht: (2025)
von: Sun, Peiwen, et al.
Veröffentlicht: (2025)
Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
von: Deng, Wei, et al.
Veröffentlicht: (2026)
von: Deng, Wei, et al.
Veröffentlicht: (2026)
pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2026)
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2026)
From Web to Pixels: Bringing Agentic Search into Visual Perception
von: Yang, Bokang, et al.
Veröffentlicht: (2026)
von: Yang, Bokang, et al.
Veröffentlicht: (2026)
RadFabric: Agentic AI System with Reasoning Capability for Radiology
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2025)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
von: Jain, Jitesh, et al.
Veröffentlicht: (2025)
von: Jain, Jitesh, et al.
Veröffentlicht: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ding, Shengyuan, et al.
Veröffentlicht: (2025)
ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning
von: Liu, Shifeng, et al.
Veröffentlicht: (2026)
von: Liu, Shifeng, et al.
Veröffentlicht: (2026)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
Act2See: Emergent Active Visual Perception for Video Reasoning
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
Self-Evolving Visual Concept Library using Vision-Language Critics
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Visual-Semantic Graph Matching Net for Zero-Shot Learning
von: Duan, Bowen, et al.
Veröffentlicht: (2024)
von: Duan, Bowen, et al.
Veröffentlicht: (2024)
NitroGen: An Open Foundation Model for Generalist Gaming Agents
von: Magne, Loïc, et al.
Veröffentlicht: (2026)
von: Magne, Loïc, et al.
Veröffentlicht: (2026)
Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
von: Yoon, Lauren Hyoseo, et al.
Veröffentlicht: (2025)
von: Yoon, Lauren Hyoseo, et al.
Veröffentlicht: (2025)
Route, Retrieve, Reflect, Repair: Self-Improving Agentic Framework for Visual Detection and Linguistic Reasoning in Medical Imaging
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2026)
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2026)
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
von: Tan, Chuangchuang, et al.
Veröffentlicht: (2025)
von: Tan, Chuangchuang, et al.
Veröffentlicht: (2025)
SARAH: Spatially Aware Real-time Agentic Humans
von: Ng, Evonne, et al.
Veröffentlicht: (2026)
von: Ng, Evonne, et al.
Veröffentlicht: (2026)
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
von: Zhao, Pukun, et al.
Veröffentlicht: (2025)
von: Zhao, Pukun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
von: Marsili, Damiano, et al.
Veröffentlicht: (2025) -
Same or Not? Enhancing Visual Perception in Vision-Language Models
von: Marsili, Damiano, et al.
Veröffentlicht: (2025) -
Find Any Part in 3D
von: Ma, Ziqi, et al.
Veröffentlicht: (2024) -
Feedforward 3D Editing via Text-Steerable Image-to-3D
von: Ma, Ziqi, et al.
Veröffentlicht: (2025) -
Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
von: Demler, Ilona, et al.
Veröffentlicht: (2025)