MResT: Multi-Resolution Sensing for Real-Time Control with Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Saxena, Saumya, Sharma, Mohit, Kroemer, Oliver |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
by: Saxena, Saumya, et al.
Published: (2024)
by: Saxena, Saumya, et al.
Published: (2024)
Test-Time Training for Visual Foresight Vision-Language-Action Models
by: Park, Sangwu, et al.
Published: (2026)
by: Park, Sangwu, et al.
Published: (2026)
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
by: Kapelyukh, Ivan, et al.
Published: (2023)
by: Kapelyukh, Ivan, et al.
Published: (2023)
Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning
by: Ding, Runyu, et al.
Published: (2024)
by: Ding, Runyu, et al.
Published: (2024)
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
by: Zhou, Xunlan, et al.
Published: (2026)
by: Zhou, Xunlan, et al.
Published: (2026)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
Tactile Modality Fusion for Vision-Language-Action Models
by: Morissette, Charlotte, et al.
Published: (2026)
by: Morissette, Charlotte, et al.
Published: (2026)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
by: Guruprasad, Pranav, et al.
Published: (2024)
by: Guruprasad, Pranav, et al.
Published: (2024)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
by: Zhang, Zezhou, et al.
Published: (2026)
by: Zhang, Zezhou, et al.
Published: (2026)
Improving Vision-Language-Action Model with Online Reinforcement Learning
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
GraphSeg: Segmented 3D Representations via Graph Edge Addition and Contraction
by: Tang, Haozhan, et al.
Published: (2025)
by: Tang, Haozhan, et al.
Published: (2025)
Semantically Controllable Augmentations for Generalizable Robot Learning
by: Chen, Zoey, et al.
Published: (2024)
by: Chen, Zoey, et al.
Published: (2024)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
by: Lee, Seungjae, et al.
Published: (2025)
by: Lee, Seungjae, et al.
Published: (2025)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
by: Li, Chengmeng, et al.
Published: (2025)
by: Li, Chengmeng, et al.
Published: (2025)
ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
by: Zhao, Yanpeng, et al.
Published: (2026)
by: Zhao, Yanpeng, et al.
Published: (2026)
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
by: Kawaharazuka, Kento, et al.
Published: (2025)
by: Kawaharazuka, Kento, et al.
Published: (2025)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
by: Shiba, Takuya
Published: (2026)
by: Shiba, Takuya
Published: (2026)
FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models
by: An, Xinyuan, et al.
Published: (2026)
by: An, Xinyuan, et al.
Published: (2026)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
Neural-Augmented Kelvinlet for Real-Time Soft Tissue Deformation Modeling
by: Shahbazi, Ashkan, et al.
Published: (2025)
by: Shahbazi, Ashkan, et al.
Published: (2025)
Vision-based Multi-future Trajectory Prediction: A Survey
by: Huang, Renhao, et al.
Published: (2023)
by: Huang, Renhao, et al.
Published: (2023)
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
by: Chahe, Amirhosein, et al.
Published: (2025)
by: Chahe, Amirhosein, et al.
Published: (2025)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024)
by: Niu, Dantong, et al.
Published: (2024)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
by: Lin, Haitao, et al.
Published: (2026)
by: Lin, Haitao, et al.
Published: (2026)
View Invariant Learning for Vision-Language Navigation in Continuous Environments
by: Sun, Josh Qixuan, et al.
Published: (2025)
by: Sun, Josh Qixuan, et al.
Published: (2025)
Continuous Object State Recognition for Cooking Robots Using Pre-Trained Vision-Language Models and Black-box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
Bridging Human Oversight and Black-box Driver Assistance: Vision-Language Models for Predictive Alerting in Lane Keeping Assist Systems
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
by: Ze, Yanjie, et al.
Published: (2023)
by: Ze, Yanjie, et al.
Published: (2023)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
by: Wang, Lirui, et al.
Published: (2023)
by: Wang, Lirui, et al.
Published: (2023)
Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation
by: Pan, Yiyuan, et al.
Published: (2024)
by: Pan, Yiyuan, et al.
Published: (2024)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
by: Xiao, Lei, et al.
Published: (2025)
by: Xiao, Lei, et al.
Published: (2025)
RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
by: Ghosh, Anurag, et al.
Published: (2026)
by: Ghosh, Anurag, et al.
Published: (2026)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
by: Xu, Siyu, et al.
Published: (2025)
by: Xu, Siyu, et al.
Published: (2025)
Hybrid Training for Vision-Language-Action Models
by: Mazzaglia, Pietro, et al.
Published: (2025)
by: Mazzaglia, Pietro, et al.
Published: (2025)
ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging
by: Iz, Selim Ahmet, et al.
Published: (2026)
by: Iz, Selim Ahmet, et al.
Published: (2026)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
by: Liang, Zhixuan, et al.
Published: (2025)
by: Liang, Zhixuan, et al.
Published: (2025)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
by: Luo, Hao, et al.
Published: (2025)
by: Luo, Hao, et al.
Published: (2025)
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
by: Vasa, Santosh, et al.
Published: (2025)
by: Vasa, Santosh, et al.
Published: (2025)
Similar Items
-
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
by: Saxena, Saumya, et al.
Published: (2024) -
Test-Time Training for Visual Foresight Vision-Language-Action Models
by: Park, Sangwu, et al.
Published: (2026) -
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
by: Kapelyukh, Ivan, et al.
Published: (2023) -
Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning
by: Ding, Runyu, et al.
Published: (2024) -
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
by: Zhou, Xunlan, et al.
Published: (2026)