SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qi, Zekun, Zhang, Wenyao, Ding, Yufei, Dong, Runpei, Yu, Xinqiang, Li, Jingwen, Xu, Lingyun, Li, Baoyu, He, Xialin, Fan, Guofan, Zhang, Jiazhao, He, Jiawei, Gu, Jiayuan, Jin, Xin, Ma, Kaisheng, Zhang, Zhizheng, Wang, He, Yi, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
von: Dong, Runpei, et al.
Veröffentlicht: (2026)
von: Dong, Runpei, et al.
Veröffentlicht: (2026)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
DexVLG: Dexterous Vision-Language-Grasp Model at Scale
von: He, Jiawei, et al.
Veröffentlicht: (2025)
von: He, Jiawei, et al.
Veröffentlicht: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
Learning Getting-Up Policies for Real-World Humanoid Robots
von: He, Xialin, et al.
Veröffentlicht: (2025)
von: He, Xialin, et al.
Veröffentlicht: (2025)
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
von: He, Xialin, et al.
Veröffentlicht: (2026)
von: He, Xialin, et al.
Veröffentlicht: (2026)
Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
Reasoning in Space via Grounding in the World
von: Chen, Yiming, et al.
Veröffentlicht: (2025)
von: Chen, Yiming, et al.
Veröffentlicht: (2025)
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
von: Qi, Zekun, et al.
Veröffentlicht: (2024)
von: Qi, Zekun, et al.
Veröffentlicht: (2024)
NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation
von: Liu, Jiahang, et al.
Veröffentlicht: (2026)
von: Liu, Jiahang, et al.
Veröffentlicht: (2026)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
von: Li, Anqi, et al.
Veröffentlicht: (2025)
von: Li, Anqi, et al.
Veröffentlicht: (2025)
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
von: Zhang, Junyu, et al.
Veröffentlicht: (2025)
von: Zhang, Junyu, et al.
Veröffentlicht: (2025)
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation
von: Liu, Jiahang, et al.
Veröffentlicht: (2026)
von: Liu, Jiahang, et al.
Veröffentlicht: (2026)
TrackVLA: Embodied Visual Tracking in the Wild
von: Wang, Shaoan, et al.
Veröffentlicht: (2025)
von: Wang, Shaoan, et al.
Veröffentlicht: (2025)
AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation
von: Zhang, Kaifeng, et al.
Veröffentlicht: (2024)
von: Zhang, Kaifeng, et al.
Veröffentlicht: (2024)
Visual Manipulation with Legs
von: He, Xialin, et al.
Veröffentlicht: (2024)
von: He, Xialin, et al.
Veröffentlicht: (2024)
Positional Prompt Tuning for Efficient 3D Representation Learning
von: Zhang, Shaochen, et al.
Veröffentlicht: (2024)
von: Zhang, Shaochen, et al.
Veröffentlicht: (2024)
Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining
von: Zhang, Wenyao, et al.
Veröffentlicht: (2026)
von: Zhang, Wenyao, et al.
Veröffentlicht: (2026)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
von: Cui, Wenbo, et al.
Veröffentlicht: (2025)
von: Cui, Wenbo, et al.
Veröffentlicht: (2025)
GAPartManip: A Large-scale Part-centric Dataset for Material-Agnostic Articulated Object Manipulation
von: Cui, Wenbo, et al.
Veröffentlicht: (2024)
von: Cui, Wenbo, et al.
Veröffentlicht: (2024)
Chromosome-level genome assembly of the Suminoe oyster Crassostrea ariakensis in south China.
von: Li, Ao, et al.
Veröffentlicht: (2024)
von: Li, Ao, et al.
Veröffentlicht: (2024)
MaskClustering: View Consensus based Mask Graph Clustering for Open-Vocabulary 3D Instance Segmentation
von: Yan, Mi, et al.
Veröffentlicht: (2024)
von: Yan, Mi, et al.
Veröffentlicht: (2024)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
DexGraspNet 2.0: Learning Generative Dexterous Grasping in Large-scale Synthetic Cluttered Scenes
von: Zhang, Jialiang, et al.
Veröffentlicht: (2024)
von: Zhang, Jialiang, et al.
Veröffentlicht: (2024)
Fostering Conservation Leadership Among the Youth: Insights From the Roots & Shoots Next Jane Program in China
von: Zhang Zekun, et al.
Veröffentlicht: (2025)
von: Zhang Zekun, et al.
Veröffentlicht: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
Composition-Grounded Data Synthesis for Visual Reasoning
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
Field Inversion Symbolic Regression with Embedded Equation Learner for Interpretable Turbulence Model Correction
von: Jiazhe, Li, et al.
Veröffentlicht: (2026)
von: Jiazhe, Li, et al.
Veröffentlicht: (2026)
Higher-Order Automatic Differentiation Using Symbolic Differential Algebra: Bridging the Gap between Algorithmic and Symbolic Differentiation
von: Zhang, He
Veröffentlicht: (2025)
von: Zhang, He
Veröffentlicht: (2025)
Bacanora and Sotol: So Far, So Close
von: Alfonso A. Gardea
Veröffentlicht: (2012)
von: Alfonso A. Gardea
Veröffentlicht: (2012)
On finite totally k-closed groups
von: He, Jiawei, et al.
Veröffentlicht: (2024)
von: He, Jiawei, et al.
Veröffentlicht: (2024)
Permutation groups and symmetric Hecke algebras
von: He, Jiawei, et al.
Veröffentlicht: (2026)
von: He, Jiawei, et al.
Veröffentlicht: (2026)
The endomorphism rings of permutation modules of $\frac{3}{2}$-transitive permutation groups
von: He, Jiawei, et al.
Veröffentlicht: (2024)
von: He, Jiawei, et al.
Veröffentlicht: (2024)
Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding
von: Zhang, Xiaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaojie, et al.
Veröffentlicht: (2025)
EfficientIML: Efficient High-Resolution Image Manipulation Localization
von: Li, Jinhan, et al.
Veröffentlicht: (2025)
von: Li, Jinhan, et al.
Veröffentlicht: (2025)
Prediction of Bridge Structural Response Based on Nonstationary Transformer
von: Qing Li, et al.
Veröffentlicht: (2025)
von: Qing Li, et al.
Veröffentlicht: (2025)
A General Theory for Compositional Generalization
von: Fu, Jingwen, et al.
Veröffentlicht: (2024)
von: Fu, Jingwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025) -
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
von: Dong, Runpei, et al.
Veröffentlicht: (2026) -
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
von: Xu, Tianyu, et al.
Veröffentlicht: (2025) -
DexVLG: Dexterous Vision-Language-Grasp Model at Scale
von: He, Jiawei, et al.
Veröffentlicht: (2025) -
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)