IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jongwoo, Ranasinghe, Kanchana, Jang, Jinhyeok, Mata, Cristina, Jang, Yoo Sung, Ryoo, Michael S |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LACE: Latent Visual Representation for Cross-Embodiment Learning
by: Jang, Yoo Sung, et al.
Published: (2026)
by: Jang, Yoo Sung, et al.
Published: (2026)
Pixel Motion as Universal Representation for Robot Control
by: Ranasinghe, Kanchana, et al.
Published: (2025)
by: Ranasinghe, Kanchana, et al.
Published: (2025)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Language Repository for Long Video Understanding
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Pixel Motion Diffusion is What We Need for Robot Control
by: Nguyen, E-Ro, et al.
Published: (2025)
by: Nguyen, E-Ro, et al.
Published: (2025)
CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
by: Mata, Cristina, et al.
Published: (2025)
by: Mata, Cristina, et al.
Published: (2025)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
by: Han, ByungOk, et al.
Published: (2024)
by: Han, ByungOk, et al.
Published: (2024)
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
DiSPo: Diffusion-SSM based Policy Learning for Coarse-to-Fine Action Discretization
by: Oh, Nayoung, et al.
Published: (2024)
by: Oh, Nayoung, et al.
Published: (2024)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
by: Park, Jongwoo, et al.
Published: (2024)
by: Park, Jongwoo, et al.
Published: (2024)
Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
by: Ye, Yifan, et al.
Published: (2025)
by: Ye, Yifan, et al.
Published: (2025)
Enhancing Diffusion Policy with Classifier-Free Guidance for Temporal Robotic Tasks
by: Lu, Yuang, et al.
Published: (2025)
by: Lu, Yuang, et al.
Published: (2025)
DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
by: Li, Yuquan, et al.
Published: (2026)
by: Li, Yuquan, et al.
Published: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
by: Peng, Xiongfeng, et al.
Published: (2026)
by: Peng, Xiongfeng, et al.
Published: (2026)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
by: Fang, Yu, et al.
Published: (2025)
by: Fang, Yu, et al.
Published: (2025)
Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
by: Neary, Cyrus, et al.
Published: (2025)
by: Neary, Cyrus, et al.
Published: (2025)
Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models
by: Lee, Jimin, et al.
Published: (2026)
by: Lee, Jimin, et al.
Published: (2026)
RoboRouter: Training-Free Policy Routing for Robotic Manipulation
by: Chen, Yiteng, et al.
Published: (2026)
by: Chen, Yiteng, et al.
Published: (2026)
EgoAVFlow: Robot Policy Learning with Active Vision from Human Egocentric Videos via 3D Flow
by: Cho, Daesol, et al.
Published: (2026)
by: Cho, Daesol, et al.
Published: (2026)
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
by: Ge, Zirui, et al.
Published: (2026)
by: Ge, Zirui, et al.
Published: (2026)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
When Attention Betrays: Erasing Backdoor Attacks in Robotic Policies by Reconstructing Visual Tokens
by: Li, Xuetao, et al.
Published: (2026)
by: Li, Xuetao, et al.
Published: (2026)
SPACE: A Python-based Simulator for Evaluating Decentralized Multi-Robot Task Allocation Algorithms
by: Jang, Inmo
Published: (2024)
by: Jang, Inmo
Published: (2024)
Social Zone as a Barrier Function for Socially-Compliant Robot Navigation
by: Jang, Junwoo, et al.
Published: (2024)
by: Jang, Junwoo, et al.
Published: (2024)
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
by: Song, Yunzhou, et al.
Published: (2026)
by: Song, Yunzhou, et al.
Published: (2026)
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
by: Kim, Dongyoung, et al.
Published: (2025)
by: Kim, Dongyoung, et al.
Published: (2025)
Functional Fibers in Soft Robotics: Advances in Material, Structural, and Systemic Tactics
by: Joonhee Won, et al.
Published: (2026)
by: Joonhee Won, et al.
Published: (2026)
CSC-MPPI: A Novel Constrained MPPI Framework with DBSCAN for Reliable Obstacle Avoidance
by: Park, Leesai, et al.
Published: (2025)
by: Park, Leesai, et al.
Published: (2025)
ILCL: Inverse Logic-Constraint Learning from Temporally Constrained Demonstrations
by: Cho, Minwoo, et al.
Published: (2025)
by: Cho, Minwoo, et al.
Published: (2025)
A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics
by: Fateh, Fawad Javed, et al.
Published: (2026)
by: Fateh, Fawad Javed, et al.
Published: (2026)
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
RA-DP: Rapid Adaptive Diffusion Policy for Training-Free High-frequency Robotics Replanning
by: Ye, Xi, et al.
Published: (2025)
by: Ye, Xi, et al.
Published: (2025)
Action-Free Reasoning for Policy Generalization
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
Understanding Long Videos with Multimodal Language Models
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
Task-Aware Positioning for Improvisational Tasks in Mobile Construction Robots via an AI Agent with Multi-LMM Modules
by: Jang, Seongju, et al.
Published: (2026)
by: Jang, Seongju, et al.
Published: (2026)
KineSoft: Learning Proprioceptive Manipulation Policies with Soft Robot Hands
by: Yoo, Uksang, et al.
Published: (2025)
by: Yoo, Uksang, et al.
Published: (2025)
GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
by: Gao, Minquan, et al.
Published: (2025)
by: Gao, Minquan, et al.
Published: (2025)
FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
by: Reuss, Moritz, et al.
Published: (2025)
by: Reuss, Moritz, et al.
Published: (2025)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
by: Kim, Seungku, et al.
Published: (2026)
by: Kim, Seungku, et al.
Published: (2026)
Similar Items
-
LACE: Latent Visual Representation for Cross-Embodiment Learning
by: Jang, Yoo Sung, et al.
Published: (2026) -
Pixel Motion as Universal Representation for Robot Control
by: Ranasinghe, Kanchana, et al.
Published: (2025) -
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
by: Li, Xiang, et al.
Published: (2024) -
Language Repository for Long Video Understanding
by: Kahatapitiya, Kumara, et al.
Published: (2024) -
Pixel Motion Diffusion is What We Need for Robot Control
by: Nguyen, E-Ro, et al.
Published: (2025)