Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Kachaev, Nikita, Kolosov, Mikhail, Zelezetsky, Daniil, Kovalev, Alexey K., Panov, Aleksandr I. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2025)
by: Cherepanov, Egor, et al.
Published: (2025)
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2026)
by: Cherepanov, Egor, et al.
Published: (2026)
Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding
by: Patratskiy, Maxim A., et al.
Published: (2025)
by: Patratskiy, Maxim A., et al.
Published: (2025)
Mind and Motion Aligned: A Joint Evaluation IsaacSim Benchmark for Task Planning and Low-Level Policies in Mobile Manipulation
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Re:Frame -- Retrieving Experience From Associative Memory
by: Zelezetsky, Daniil, et al.
Published: (2025)
by: Zelezetsky, Daniil, et al.
Published: (2025)
VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots
by: Grigorev, Danil S., et al.
Published: (2025)
by: Grigorev, Danil S., et al.
Published: (2025)
ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
by: Cherepanov, Egor, et al.
Published: (2025)
by: Cherepanov, Egor, et al.
Published: (2025)
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation
by: Cherepanov, Egor, et al.
Published: (2024)
by: Cherepanov, Egor, et al.
Published: (2024)
LookPlanGraph: Embodied Instruction Following Method with VLM Graph Augmentation
by: Onishchenko, Anatoly O., et al.
Published: (2025)
by: Onishchenko, Anatoly O., et al.
Published: (2025)
Accelerating Transformers in Online RL
by: Zelezetsky, Daniil, et al.
Published: (2025)
by: Zelezetsky, Daniil, et al.
Published: (2025)
AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment
by: Ivanova, Anastasiia, et al.
Published: (2025)
by: Ivanova, Anastasiia, et al.
Published: (2025)
HELP: Hierarchical Embodied Language Planner for Household Tasks
by: Korchemnyi, Alexandr V., et al.
Published: (2025)
by: Korchemnyi, Alexandr V., et al.
Published: (2025)
Symbolic Disentangled Representations for Images
by: Korchemnyi, Alexandr, et al.
Published: (2024)
by: Korchemnyi, Alexandr, et al.
Published: (2024)
Object-Centric Learning with Slot Mixture Module
by: Kirilenko, Daniil, et al.
Published: (2023)
by: Kirilenko, Daniil, et al.
Published: (2023)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning
by: Nesterova, Maria, et al.
Published: (2026)
by: Nesterova, Maria, et al.
Published: (2026)
ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
by: Zhang, Zhuoyang, et al.
Published: (2026)
by: Zhang, Zhuoyang, et al.
Published: (2026)
Recurrent Action Transformer with Memory
by: Cherepanov, Egor, et al.
Published: (2023)
by: Cherepanov, Egor, et al.
Published: (2023)
Relational Object-Centric Actor-Critic
by: Ugadiarov, Leonid, et al.
Published: (2023)
by: Ugadiarov, Leonid, et al.
Published: (2023)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
by: Zhang, Jiefu, et al.
Published: (2026)
by: Zhang, Jiefu, et al.
Published: (2026)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
Object-Centric World Models Meet Monte Carlo Tree Search
by: Vakhitov, Rodion, et al.
Published: (2026)
by: Vakhitov, Rodion, et al.
Published: (2026)
"Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generation
by: Seffo, Amin, et al.
Published: (2025)
by: Seffo, Amin, et al.
Published: (2025)
Memory Retention Is Not Enough to Master Memory Tasks in Reinforcement Learning
by: Shchendrigin, Oleg, et al.
Published: (2026)
by: Shchendrigin, Oleg, et al.
Published: (2026)
Dynamic Neural Potential Field: Online Trajectory Optimization in the Presence of Moving Obstacles
by: Staroverov, Aleksei, et al.
Published: (2024)
by: Staroverov, Aleksei, et al.
Published: (2024)
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
by: Li, Haoyun, et al.
Published: (2025)
by: Li, Haoyun, et al.
Published: (2025)
Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting
by: Schaldenbrand, Peter, et al.
Published: (2026)
by: Schaldenbrand, Peter, et al.
Published: (2026)
Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA
by: Mai, Hung, et al.
Published: (2026)
by: Mai, Hung, et al.
Published: (2026)
Knowledge-Guided Manipulation Using Multi-Task Reinforcement Learning
by: Narendra, Aditya, et al.
Published: (2026)
by: Narendra, Aditya, et al.
Published: (2026)
SmoothVLA: Aligning Vision-Language-Action Models with Physical Constraints via Intrinsic Smoothness Optimization
by: Li, Jiashun, et al.
Published: (2026)
by: Li, Jiashun, et al.
Published: (2026)
DroneVLA: VLA based Aerial Manipulation
by: Mehboob, Fawad, et al.
Published: (2026)
by: Mehboob, Fawad, et al.
Published: (2026)
Atomic Action Slicing: Planner-Aligned Options for Generalist VLA Agents
by: Tabakov, Stefan, et al.
Published: (2025)
by: Tabakov, Stefan, et al.
Published: (2025)
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
Don't Start from Scratch: Behavioral Refinement via Interpolant-based Policy Diffusion
by: Chen, Kaiqi, et al.
Published: (2024)
by: Chen, Kaiqi, et al.
Published: (2024)
LERa: Replanning with Visual Feedback in Instruction Following
by: Pchelintsev, Svyatoslav, et al.
Published: (2025)
by: Pchelintsev, Svyatoslav, et al.
Published: (2025)
RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour
by: Serpiva, Valerii, et al.
Published: (2025)
by: Serpiva, Valerii, et al.
Published: (2025)
TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
by: Zhong, Yiming, et al.
Published: (2026)
by: Zhong, Yiming, et al.
Published: (2026)
Sci-VLA: Agentic VLA Inference Plugin for Long-Horizon Tasks in Scientific Experiments
by: Pang, Yiwen, et al.
Published: (2026)
by: Pang, Yiwen, et al.
Published: (2026)
Similar Items
-
A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
by: Kachaev, Nikita, et al.
Published: (2025) -
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2025) -
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2026) -
Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding
by: Patratskiy, Maxim A., et al.
Published: (2025) -
Mind and Motion Aligned: A Joint Evaluation IsaacSim Benchmark for Task Planning and Low-Level Policies in Mobile Manipulation
by: Kachaev, Nikita, et al.
Published: (2025)