Grounding Intelligence in Movement
Fuente:
arXiv
Guardado en:
| Autores principales: | Segado, Melanie, Parodi, Felipe, Matelsky, Jordan K., Platt, Michael L., Dyer, Eva B., Kording, Konrad P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers
por: Parodi, Felipe, et al.
Publicado: (2026)
por: Parodi, Felipe, et al.
Publicado: (2026)
Vision-language models for decoding provider attention during neonatal resuscitation
por: Parodi, Felipe, et al.
Publicado: (2024)
por: Parodi, Felipe, et al.
Publicado: (2024)
Equivariant Reinforcement Learning under Partial Observability
por: Nguyen, Hai, et al.
Publicado: (2024)
por: Nguyen, Hai, et al.
Publicado: (2024)
Empirical influence functions to understand the logic of fine-tuning
por: Matelsky, Jordan K., et al.
Publicado: (2024)
por: Matelsky, Jordan K., et al.
Publicado: (2024)
Grounding Driving VLA via Inverse Kinematics
por: Park, Junsung, et al.
Publicado: (2026)
por: Park, Junsung, et al.
Publicado: (2026)
Physically Grounded Vision-Language Models for Robotic Manipulation
por: Gao, Jensen, et al.
Publicado: (2023)
por: Gao, Jensen, et al.
Publicado: (2023)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
por: Xiao, Junbin, et al.
Publicado: (2026)
por: Xiao, Junbin, et al.
Publicado: (2026)
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
por: Zhou, Yang, et al.
Publicado: (2026)
por: Zhou, Yang, et al.
Publicado: (2026)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
por: Chen, Shizhe, et al.
Publicado: (2025)
por: Chen, Shizhe, et al.
Publicado: (2025)
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
por: Saxena, Pranav, et al.
Publicado: (2025)
por: Saxena, Pranav, et al.
Publicado: (2025)
ContactGaussian-WM: Learning Physics-Grounded World Model from Videos
por: Wang, Meizhong, et al.
Publicado: (2026)
por: Wang, Meizhong, et al.
Publicado: (2026)
AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration
por: Gao, Xiangbo, et al.
Publicado: (2025)
por: Gao, Xiangbo, et al.
Publicado: (2025)
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
por: Qi, Zekun, et al.
Publicado: (2025)
por: Qi, Zekun, et al.
Publicado: (2025)
Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
por: Yu, Seungjun, et al.
Publicado: (2025)
por: Yu, Seungjun, et al.
Publicado: (2025)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
por: Zhu, He, et al.
Publicado: (2025)
por: Zhu, He, et al.
Publicado: (2025)
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
por: Yang, Tianshuo, et al.
Publicado: (2026)
por: Yang, Tianshuo, et al.
Publicado: (2026)
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
por: Zhang, Ninghao, et al.
Publicado: (2026)
por: Zhang, Ninghao, et al.
Publicado: (2026)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
por: Feng, ZhiYuan, et al.
Publicado: (2026)
por: Feng, ZhiYuan, et al.
Publicado: (2026)
On the Application of Efficient Neural Mapping to Real-Time Indoor Localisation for Unmanned Ground Vehicles
por: Holder, Christopher J., et al.
Publicado: (2022)
por: Holder, Christopher J., et al.
Publicado: (2022)
Vision-Language Navigation with Embodied Intelligence: A Survey
por: Gao, Peng, et al.
Publicado: (2024)
por: Gao, Peng, et al.
Publicado: (2024)
Dejavu: Towards Experience Feedback Learning for Embodied Intelligence
por: Wu, Shaokai, et al.
Publicado: (2025)
por: Wu, Shaokai, et al.
Publicado: (2025)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
por: Lyu, Ruiyuan, et al.
Publicado: (2024)
por: Lyu, Ruiyuan, et al.
Publicado: (2024)
GVDepth: Zero-Shot Monocular Depth Estimation for Ground Vehicles based on Probabilistic Cue Fusion
por: Koledić, Karlo, et al.
Publicado: (2024)
por: Koledić, Karlo, et al.
Publicado: (2024)
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
por: Jahangard, Simindokht, et al.
Publicado: (2025)
por: Jahangard, Simindokht, et al.
Publicado: (2025)
Kick Back & Relax++: Scaling Beyond Ground-Truth Depth with SlowTV & CribsTV
por: Spencer, Jaime, et al.
Publicado: (2024)
por: Spencer, Jaime, et al.
Publicado: (2024)
Dynamic Open Vocabulary Enhanced Safe-landing with Intelligence (DOVESEI)
por: Bong, Haechan Mark, et al.
Publicado: (2023)
por: Bong, Haechan Mark, et al.
Publicado: (2023)
FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation
por: Zeng, Huajian, et al.
Publicado: (2026)
por: Zeng, Huajian, et al.
Publicado: (2026)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
por: Fang, Jiading
Publicado: (2025)
por: Fang, Jiading
Publicado: (2025)
OMEGA: Efficient Occlusion-Aware Navigation for Air-Ground Robot in Dynamic Environments via State Space Model
por: Wang, Junming, et al.
Publicado: (2024)
por: Wang, Junming, et al.
Publicado: (2024)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
por: Zhang, Pingrui, et al.
Publicado: (2025)
por: Zhang, Pingrui, et al.
Publicado: (2025)
UAV-based Intelligent Information Systems on Winter Road Safety for Autonomous Vehicles
por: Ariram, Siva, et al.
Publicado: (2024)
por: Ariram, Siva, et al.
Publicado: (2024)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
por: Zantout, Nader, et al.
Publicado: (2025)
por: Zantout, Nader, et al.
Publicado: (2025)
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
por: Ci, Hai, et al.
Publicado: (2025)
por: Ci, Hai, et al.
Publicado: (2025)
Go-SLAM: Grounded Object Segmentation and Localization with Gaussian Splatting SLAM
por: Pham, Phu, et al.
Publicado: (2024)
por: Pham, Phu, et al.
Publicado: (2024)
A Systematic Literature Review of Computer Vision Applications in Robotized Wire Harness Assembly
por: Wang, Hao, et al.
Publicado: (2023)
por: Wang, Hao, et al.
Publicado: (2023)
AnyTraverse: An off-road traversability framework with VLM and human operator in the loop
por: Sahu, Sattwik, et al.
Publicado: (2025)
por: Sahu, Sattwik, et al.
Publicado: (2025)
CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence
por: Zeng, Tianle, et al.
Publicado: (2026)
por: Zeng, Tianle, et al.
Publicado: (2026)
Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
por: Li, Yihao, et al.
Publicado: (2025)
por: Li, Yihao, et al.
Publicado: (2025)
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation
por: Govindarajan, Hariprasath, et al.
Publicado: (2025)
por: Govindarajan, Hariprasath, et al.
Publicado: (2025)
S3PT: Scene Semantics and Structure Guided Clustering to Boost Self-Supervised Pre-Training for Autonomous Driving
por: Wozniak, Maciej K., et al.
Publicado: (2024)
por: Wozniak, Maciej K., et al.
Publicado: (2024)
Ejemplares similares
-
Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers
por: Parodi, Felipe, et al.
Publicado: (2026) -
Vision-language models for decoding provider attention during neonatal resuscitation
por: Parodi, Felipe, et al.
Publicado: (2024) -
Equivariant Reinforcement Learning under Partial Observability
por: Nguyen, Hai, et al.
Publicado: (2024) -
Empirical influence functions to understand the logic of fine-tuning
por: Matelsky, Jordan K., et al.
Publicado: (2024) -
Grounding Driving VLA via Inverse Kinematics
por: Park, Junsung, et al.
Publicado: (2026)