OmniSAT: Compact Action Token, Faster Auto Regression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lyu, Huaihai, Chen, Chaofan, Xie, Senwei, Wang, Pengwei, Chen, Xiansheng, Zhang, Shanghang, Xu, Changsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
von: Lyu, Huaihai, et al.
Veröffentlicht: (2026)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2026)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
von: Fu, Yankai, et al.
Veröffentlicht: (2025)
von: Fu, Yankai, et al.
Veröffentlicht: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
OmniIndoor3D: Comprehensive Indoor 3D Reconstruction
von: Wei, Xiaobao, et al.
Veröffentlicht: (2025)
von: Wei, Xiaobao, et al.
Veröffentlicht: (2025)
PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing
von: Ji, Yuheng, et al.
Veröffentlicht: (2026)
von: Ji, Yuheng, et al.
Veröffentlicht: (2026)
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
von: Chen, Anthony, et al.
Veröffentlicht: (2025)
von: Chen, Anthony, et al.
Veröffentlicht: (2025)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
von: Wang, Zifan, et al.
Veröffentlicht: (2024)
von: Wang, Zifan, et al.
Veröffentlicht: (2024)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
von: Hao, Peng, et al.
Veröffentlicht: (2025)
von: Hao, Peng, et al.
Veröffentlicht: (2025)
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
von: Peng, Cheng, et al.
Veröffentlicht: (2025)
von: Peng, Cheng, et al.
Veröffentlicht: (2025)
Target-Oriented Object Grasping via Multimodal Human Guidance
von: Xie, Pengwei, et al.
Veröffentlicht: (2024)
von: Xie, Pengwei, et al.
Veröffentlicht: (2024)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
Enhancing Scene Coordinate Regression with Efficient Keypoint Detection and Sequential Information
von: Xu, Kuan, et al.
Veröffentlicht: (2024)
von: Xu, Kuan, et al.
Veröffentlicht: (2024)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
von: Li, Zaijing, et al.
Veröffentlicht: (2026)
von: Li, Zaijing, et al.
Veröffentlicht: (2026)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
von: Li, Haosheng, et al.
Veröffentlicht: (2026)
von: Li, Haosheng, et al.
Veröffentlicht: (2026)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
W-PoseNet: Dense Correspondence Regularized Pixel Pair Pose Regression
von: Xu, Zelin, et al.
Veröffentlicht: (2019)
von: Xu, Zelin, et al.
Veröffentlicht: (2019)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
Efficient Heatmap-Guided 6-Dof Grasp Detection in Cluttered Scenes
von: Chen, Siang, et al.
Veröffentlicht: (2024)
von: Chen, Siang, et al.
Veröffentlicht: (2024)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
von: Han, Yi, et al.
Veröffentlicht: (2025)
von: Han, Yi, et al.
Veröffentlicht: (2025)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
von: Zeng, Zichao, et al.
Veröffentlicht: (2026)
von: Zeng, Zichao, et al.
Veröffentlicht: (2026)
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
von: Liu, Katherine, et al.
Veröffentlicht: (2025)
von: Liu, Katherine, et al.
Veröffentlicht: (2025)
AoE: Always-on Egocentric Human Video Collection for Embodied AI
von: Yang, Bowen, et al.
Veröffentlicht: (2026)
von: Yang, Bowen, et al.
Veröffentlicht: (2026)
OmniLiDAR: A Unified Diffusion Framework for Multi-Domain 3D LiDAR Generation
von: Liu, Youquan, et al.
Veröffentlicht: (2026)
von: Liu, Youquan, et al.
Veröffentlicht: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
DriveVA: Video Action Models are Zero-Shot Drivers
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
von: Lyu, Huaihai, et al.
Veröffentlicht: (2026) -
EgoPrompt: Prompt Learning for Egocentric Action Recognition
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025) -
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
von: Fu, Yankai, et al.
Veröffentlicht: (2025) -
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026) -
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)