Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
Fuente:
arXiv
Guardado en:
| Autores principales: | Bendikas, Rokas, Dijkman, Daniel, Peschl, Markus, Haresh, Sanjay, Mazzaglia, Pietro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hybrid Training for Vision-Language-Action Models
por: Mazzaglia, Pietro, et al.
Publicado: (2025)
por: Mazzaglia, Pietro, et al.
Publicado: (2025)
From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
por: Peschl, Markus, et al.
Publicado: (2025)
por: Peschl, Markus, et al.
Publicado: (2025)
Information-driven Affordance Discovery for Efficient Robotic Manipulation
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
What Matters in Building Vision-Language-Action Models for Generalist Robots
por: Li, Xinghang, et al.
Publicado: (2024)
por: Li, Xinghang, et al.
Publicado: (2024)
Redundancy-aware Action Spaces for Robot Learning
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
por: Feng, Yixu, et al.
Publicado: (2026)
por: Feng, Yixu, et al.
Publicado: (2026)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
por: Liu, Yicheng, et al.
Publicado: (2025)
por: Liu, Yicheng, et al.
Publicado: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
por: Lin, Yihan, et al.
Publicado: (2026)
por: Lin, Yihan, et al.
Publicado: (2026)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
por: Wang, Yating, et al.
Publicado: (2025)
por: Wang, Yating, et al.
Publicado: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
por: Xie, Haozhe, et al.
Publicado: (2026)
por: Xie, Haozhe, et al.
Publicado: (2026)
Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
por: Sohn, Tin Stribor, et al.
Publicado: (2025)
por: Sohn, Tin Stribor, et al.
Publicado: (2025)
VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility
por: Shi, Yitian, et al.
Publicado: (2025)
por: Shi, Yitian, et al.
Publicado: (2025)
GenRL: Multimodal-foundation world models for generalization in embodied agents
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
por: Mazzaglia, Pietro, et al.
Publicado: (2024)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
por: Xue, Wei, et al.
Publicado: (2026)
por: Xue, Wei, et al.
Publicado: (2026)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
por: Li, Chenyang, et al.
Publicado: (2026)
por: Li, Chenyang, et al.
Publicado: (2026)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
por: Li, Haosheng, et al.
Publicado: (2026)
por: Li, Haosheng, et al.
Publicado: (2026)
Information-driven Affordance Discovery for Efficient Robotic Manipulation
por: Mazzaglia, Pietro, et al.
Publicado: (2023)
por: Mazzaglia, Pietro, et al.
Publicado: (2023)
From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models
por: Pulli, Tessa, et al.
Publicado: (2024)
por: Pulli, Tessa, et al.
Publicado: (2024)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
por: Zhou, Jiaying, et al.
Publicado: (2026)
por: Zhou, Jiaying, et al.
Publicado: (2026)
Unified Vision-Language-Action Model
por: Wang, Yuqi, et al.
Publicado: (2025)
por: Wang, Yuqi, et al.
Publicado: (2025)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
por: Shiba, Takuya
Publicado: (2026)
por: Shiba, Takuya
Publicado: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
por: Xu, Siyu, et al.
Publicado: (2025)
por: Xu, Siyu, et al.
Publicado: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
por: Li, Qiwei, et al.
Publicado: (2026)
por: Li, Qiwei, et al.
Publicado: (2026)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
por: Jeong, Seunghoon, et al.
Publicado: (2026)
por: Jeong, Seunghoon, et al.
Publicado: (2026)
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
por: Halacheva, Anna-Maria, et al.
Publicado: (2025)
por: Halacheva, Anna-Maria, et al.
Publicado: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
por: Zhang, Tianyi, et al.
Publicado: (2025)
por: Zhang, Tianyi, et al.
Publicado: (2025)
Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs
por: Orjuela, Daniel Yezid Guarnizo, et al.
Publicado: (2026)
por: Orjuela, Daniel Yezid Guarnizo, et al.
Publicado: (2026)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
por: Peng, Jierui, et al.
Publicado: (2025)
por: Peng, Jierui, et al.
Publicado: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
por: Chen, Shizhe, et al.
Publicado: (2026)
por: Chen, Shizhe, et al.
Publicado: (2026)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
por: Lv, Qi, et al.
Publicado: (2025)
por: Lv, Qi, et al.
Publicado: (2025)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
por: Zhao, Chen, et al.
Publicado: (2026)
por: Zhao, Chen, et al.
Publicado: (2026)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
por: Govind, Manish Kumar, et al.
Publicado: (2026)
por: Govind, Manish Kumar, et al.
Publicado: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
por: Ding, Pengxiang, et al.
Publicado: (2023)
por: Ding, Pengxiang, et al.
Publicado: (2023)
LLaDA-VLA: Vision Language Diffusion Action Models
por: Wen, Yuqing, et al.
Publicado: (2025)
por: Wen, Yuqing, et al.
Publicado: (2025)
That's My Point: Compact Object-centric LiDAR Pose Estimation for Large-scale Outdoor Localisation
por: Pramatarov, Georgi, et al.
Publicado: (2024)
por: Pramatarov, Georgi, et al.
Publicado: (2024)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
por: Wu, Yuhan, et al.
Publicado: (2025)
por: Wu, Yuhan, et al.
Publicado: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
por: Sarowar, Md Selim, et al.
Publicado: (2026)
por: Sarowar, Md Selim, et al.
Publicado: (2026)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
por: Song, Wenxuan, et al.
Publicado: (2025)
por: Song, Wenxuan, et al.
Publicado: (2025)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
por: Zhang, Chubin, et al.
Publicado: (2026)
por: Zhang, Chubin, et al.
Publicado: (2026)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
por: Ye, Wencheng, et al.
Publicado: (2025)
por: Ye, Wencheng, et al.
Publicado: (2025)
Ejemplares similares
-
Hybrid Training for Vision-Language-Action Models
por: Mazzaglia, Pietro, et al.
Publicado: (2025) -
From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
por: Peschl, Markus, et al.
Publicado: (2025) -
Information-driven Affordance Discovery for Efficient Robotic Manipulation
por: Mazzaglia, Pietro, et al.
Publicado: (2024) -
What Matters in Building Vision-Language-Action Models for Generalist Robots
por: Li, Xinghang, et al.
Publicado: (2024) -
Redundancy-aware Action Spaces for Robot Learning
por: Mazzaglia, Pietro, et al.
Publicado: (2024)