T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yiteng, Li, Wenbo, Wang, Shiyi, Zhuang, Huiping, Wu, Qingyao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
by: Li, Wenbo, et al.
Published: (2025)
by: Li, Wenbo, et al.
Published: (2025)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
by: Mehta, Vinit, et al.
Published: (2025)
by: Mehta, Vinit, et al.
Published: (2025)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
by: Xiao, Jiasong, et al.
Published: (2026)
by: Xiao, Jiasong, et al.
Published: (2026)
Unveiling the Potential of iMarkers: Invisible Fiducial Markers for Advanced Robotics
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
Temporally Consistent Object 6D Pose Estimation for Robot Control
by: Zorina, Kateryna, et al.
Published: (2026)
by: Zorina, Kateryna, et al.
Published: (2026)
Autonomous Underwater Cognitive System for Adaptive Navigation: A SLAM-Integrated Cognitive Architecture
by: Jayarathne, K. A. I. N, et al.
Published: (2025)
by: Jayarathne, K. A. I. N, et al.
Published: (2025)
SuperPoint-SLAM3: Augmenting ORB-SLAM3 with Deep Features, Adaptive NMS, and Learning-Based Loop Closure
by: Syed, Shahram Najam, et al.
Published: (2025)
by: Syed, Shahram Najam, et al.
Published: (2025)
vS-Graphs: Tightly Coupling Visual SLAM and 3D Scene Graphs Exploiting Hierarchical Scene Understanding
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
by: Endo, Masafumi, et al.
Published: (2024)
by: Endo, Masafumi, et al.
Published: (2024)
Is Single-View Mesh Reconstruction Ready for Robotics?
by: Nolte, Frederik, et al.
Published: (2025)
by: Nolte, Frederik, et al.
Published: (2025)
Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
by: Käs, Stephanie, et al.
Published: (2025)
by: Käs, Stephanie, et al.
Published: (2025)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
by: Hu, X., et al.
Published: (2025)
by: Hu, X., et al.
Published: (2025)
Botany Meets Robotics in Alpine Scree Monitoring
by: De Benedittis, Davide, et al.
Published: (2025)
by: De Benedittis, Davide, et al.
Published: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
by: Riva, Paolo, et al.
Published: (2026)
by: Riva, Paolo, et al.
Published: (2026)
FrankenBot: Brain-Morphic Modular Orchestration for Robotic Manipulation with Vision-Language Models
by: Wang, Shiyi, et al.
Published: (2025)
by: Wang, Shiyi, et al.
Published: (2025)
Evaluating the Impact of Synthetic Data on Object Detection Tasks in Autonomous Driving
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
Introspection in Learned Semantic Scene Graph Localisation
by: Bissessur, Manshika Charvi, et al.
Published: (2025)
by: Bissessur, Manshika Charvi, et al.
Published: (2025)
How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?
by: Käs, Stephanie, et al.
Published: (2025)
by: Käs, Stephanie, et al.
Published: (2025)
Distributed Intelligent System Architecture for UAV-Assisted Monitoring of Wind Energy Infrastructure
by: Svystun, Serhii, et al.
Published: (2024)
by: Svystun, Serhii, et al.
Published: (2024)
CARScenes: Semantic VLM Dataset for Safe Autonomous Driving
by: He, Yuankai, et al.
Published: (2025)
by: He, Yuankai, et al.
Published: (2025)
Have We Mastered Scale in Deep Monocular Visual SLAM? The ScaleMaster Dataset and Benchmark
by: Ju, Hyoseok, et al.
Published: (2026)
by: Ju, Hyoseok, et al.
Published: (2026)
CODEI: Resource-Efficient Task-Driven Co-Design of Perception and Decision Making for Mobile Robots Applied to Autonomous Vehicles
by: Milojevic, Dejan, et al.
Published: (2025)
by: Milojevic, Dejan, et al.
Published: (2025)
Immersive Robot Programming Interface for Human-Guided Automation and Randomized Path Planning
by: Malek, Kaveh, et al.
Published: (2024)
by: Malek, Kaveh, et al.
Published: (2024)
Single-Shot Metric Depth from Focused Plenoptic Cameras
by: Lasheras-Hernandez, Blanca, et al.
Published: (2024)
by: Lasheras-Hernandez, Blanca, et al.
Published: (2024)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
by: Tourani, Ali, et al.
Published: (2023)
by: Tourani, Ali, et al.
Published: (2023)
Key-Scan-Based Mobile Robot Navigation: Integrated Mapping, Planning, and Control using Graphs of Scan Regions
by: Latha, Dharshan Bashkaran, et al.
Published: (2024)
by: Latha, Dharshan Bashkaran, et al.
Published: (2024)
From Photons to Physics: Autonomous Indoor Drones and the Future of Objective Property Assessment
by: Teikari, Petteri, et al.
Published: (2025)
by: Teikari, Petteri, et al.
Published: (2025)
SafeDMPs: Integrating Formal Safety with DMPs for Adaptive HRI
by: Nath, Soumyodipta, et al.
Published: (2026)
by: Nath, Soumyodipta, et al.
Published: (2026)
Diffusion-SAFE: Diffusion-Native Human-to-Robot Driving Handover for Shared Autonomy
by: Fan, Yunxin, et al.
Published: (2025)
by: Fan, Yunxin, et al.
Published: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025)
by: Romero, Angel, et al.
Published: (2025)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
by: Bonial, Claire, et al.
Published: (2024)
by: Bonial, Claire, et al.
Published: (2024)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
by: Lukin, Stephanie M., et al.
Published: (2024)
by: Lukin, Stephanie M., et al.
Published: (2024)
M2HRI: An LLM-Driven Multimodal Multi-Agent Framework for Personalized Human-Robot Interaction
by: Hasan, Shaid, et al.
Published: (2026)
by: Hasan, Shaid, et al.
Published: (2026)
FCBV-Net: Category-Level Robotic Garment Smoothing via Feature-Conditioned Bimanual Value Prediction
by: Daba, Mohammed, et al.
Published: (2025)
by: Daba, Mohammed, et al.
Published: (2025)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
by: Römer, Ralf, et al.
Published: (2025)
by: Römer, Ralf, et al.
Published: (2025)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
by: Ferenczi, Bryce, et al.
Published: (2023)
by: Ferenczi, Bryce, et al.
Published: (2023)
Thermal and RGB Images Work Better Together in Wind Turbine Damage Detection
by: Svystun, Serhii, et al.
Published: (2024)
by: Svystun, Serhii, et al.
Published: (2024)
Failure Prediction at Runtime for Generative Robot Policies
by: Römer, Ralf, et al.
Published: (2025)
by: Römer, Ralf, et al.
Published: (2025)
Similar Items
-
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
by: Li, Wenbo, et al.
Published: (2025) -
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
by: Mehta, Vinit, et al.
Published: (2025) -
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
by: Xiao, Jiasong, et al.
Published: (2026) -
Unveiling the Potential of iMarkers: Invisible Fiducial Markers for Advanced Robotics
by: Tourani, Ali, et al.
Published: (2025) -
Temporally Consistent Object 6D Pose Estimation for Robot Control
by: Zorina, Kateryna, et al.
Published: (2026)