SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
Fuente:
arXiv
Guardado en:
| Autores principales: | Traore, Abdarahmane, Hervet, Éric, Couturier, Andy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Fungi Prototype Representations for Few-Shot Classification
por: Traore, Abdarahmane, et al.
Publicado: (2025)
por: Traore, Abdarahmane, et al.
Publicado: (2025)
Violence detection in videos using deep recurrent and convolutional neural networks
por: Traoré, Abdarahmane, et al.
Publicado: (2024)
por: Traoré, Abdarahmane, et al.
Publicado: (2024)
2D bidirectional gated recurrent unit convolutional Neural networks for end-to-end violence detection In videos
por: Traoré, Abdarahmane, et al.
Publicado: (2024)
por: Traoré, Abdarahmane, et al.
Publicado: (2024)
SmolVLM: Redefining small and efficient multimodal models
por: Marafioti, Andrés, et al.
Publicado: (2025)
por: Marafioti, Andrés, et al.
Publicado: (2025)
Warehouse Spatial Question Answering with LLM Agent
por: Huang, Hsiang-Wei, et al.
Publicado: (2025)
por: Huang, Hsiang-Wei, et al.
Publicado: (2025)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
por: Cheng, An-Chieh, et al.
Publicado: (2024)
por: Cheng, An-Chieh, et al.
Publicado: (2024)
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
por: Xu, Haotian, et al.
Publicado: (2026)
por: Xu, Haotian, et al.
Publicado: (2026)
KernelWarehouse: Rethinking the Design of Dynamic Convolution
por: Li, Chao, et al.
Publicado: (2024)
por: Li, Chao, et al.
Publicado: (2024)
Geometrically-Constrained Agent for Spatial Reasoning
por: Chen, Zeren, et al.
Publicado: (2025)
por: Chen, Zeren, et al.
Publicado: (2025)
Pursuing Minimal Sufficiency in Spatial Reasoning
por: Guo, Yejie, et al.
Publicado: (2025)
por: Guo, Yejie, et al.
Publicado: (2025)
Make Geometry Matter for Spatial Reasoning
por: Zhang, Shihua, et al.
Publicado: (2026)
por: Zhang, Shihua, et al.
Publicado: (2026)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
por: Pan, Zhenyu, et al.
Publicado: (2025)
por: Pan, Zhenyu, et al.
Publicado: (2025)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
por: Wu, Hang, et al.
Publicado: (2026)
por: Wu, Hang, et al.
Publicado: (2026)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
por: Xu, Zelin, et al.
Publicado: (2026)
por: Xu, Zelin, et al.
Publicado: (2026)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
por: Lee, Youngwan, et al.
Publicado: (2026)
por: Lee, Youngwan, et al.
Publicado: (2026)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
por: Chen, Zizhao, et al.
Publicado: (2025)
por: Chen, Zizhao, et al.
Publicado: (2025)
Enhancing Spatial Reasoning through Visual and Textual Thinking
por: Liang, Xun, et al.
Publicado: (2025)
por: Liang, Xun, et al.
Publicado: (2025)
Limits of Spatial Imagery Reasoning in Frontier LLM Models
por: Hayashi, Sergio Y., et al.
Publicado: (2026)
por: Hayashi, Sergio Y., et al.
Publicado: (2026)
CoV: Chain-of-View Prompting for Spatial Reasoning
por: Zhao, Haoyu, et al.
Publicado: (2026)
por: Zhao, Haoyu, et al.
Publicado: (2026)
Parameter-Efficient Active Learning for Foundational models
por: Narayanan, Athmanarayanan Lakshmi, et al.
Publicado: (2024)
por: Narayanan, Athmanarayanan Lakshmi, et al.
Publicado: (2024)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
por: Wang, Xingrui, et al.
Publicado: (2025)
por: Wang, Xingrui, et al.
Publicado: (2025)
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
por: Chunhachatrachai, Pawat, et al.
Publicado: (2026)
por: Chunhachatrachai, Pawat, et al.
Publicado: (2026)
Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting
por: Bhyri, Rishikesh, et al.
Publicado: (2026)
por: Bhyri, Rishikesh, et al.
Publicado: (2026)
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
por: Shiri, Fatemeh, et al.
Publicado: (2024)
por: Shiri, Fatemeh, et al.
Publicado: (2024)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
por: Ma, Xueqi, et al.
Publicado: (2026)
por: Ma, Xueqi, et al.
Publicado: (2026)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
por: Kim, Kwonyoung, et al.
Publicado: (2025)
por: Kim, Kwonyoung, et al.
Publicado: (2025)
Efficient Parameter Mining and Freezing for Continual Object Detection
por: Menezes, Angelo G., et al.
Publicado: (2024)
por: Menezes, Angelo G., et al.
Publicado: (2024)
StarCraftImage: A Dataset For Prototyping Spatial Reasoning Methods For Multi-Agent Environments
por: Kulinski, Sean, et al.
Publicado: (2024)
por: Kulinski, Sean, et al.
Publicado: (2024)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
por: Tian, Kexin, et al.
Publicado: (2025)
por: Tian, Kexin, et al.
Publicado: (2025)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
por: Liu, Zishan, et al.
Publicado: (2026)
por: Liu, Zishan, et al.
Publicado: (2026)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
por: AI, Inclusion, et al.
Publicado: (2025)
por: AI, Inclusion, et al.
Publicado: (2025)
Texo: Formula Recognition within 20M Parameters
por: Mao, Sicheng
Publicado: (2026)
por: Mao, Sicheng
Publicado: (2026)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
por: Fang, Jiading
Publicado: (2025)
por: Fang, Jiading
Publicado: (2025)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
por: Liu, Chonghan, et al.
Publicado: (2025)
por: Liu, Chonghan, et al.
Publicado: (2025)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
por: Zhang, Yuyou, et al.
Publicado: (2025)
por: Zhang, Yuyou, et al.
Publicado: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
por: Taioli, Francesco, et al.
Publicado: (2026)
por: Taioli, Francesco, et al.
Publicado: (2026)
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
por: Xue, Qiyao, et al.
Publicado: (2025)
por: Xue, Qiyao, et al.
Publicado: (2025)
Multimodal Parameter-Efficient Few-Shot Class Incremental Learning
por: D'Alessandro, Marco, et al.
Publicado: (2023)
por: D'Alessandro, Marco, et al.
Publicado: (2023)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
por: Li, Hongxing, et al.
Publicado: (2025)
por: Li, Hongxing, et al.
Publicado: (2025)
Ejemplares similares
-
Improving Fungi Prototype Representations for Few-Shot Classification
por: Traore, Abdarahmane, et al.
Publicado: (2025) -
Violence detection in videos using deep recurrent and convolutional neural networks
por: Traoré, Abdarahmane, et al.
Publicado: (2024) -
2D bidirectional gated recurrent unit convolutional Neural networks for end-to-end violence detection In videos
por: Traoré, Abdarahmane, et al.
Publicado: (2024) -
SmolVLM: Redefining small and efficient multimodal models
por: Marafioti, Andrés, et al.
Publicado: (2025) -
Warehouse Spatial Question Answering with LLM Agent
por: Huang, Hsiang-Wei, et al.
Publicado: (2025)