ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Cheng, Jiao, Jianhao, Huang, Lingyi, Xiao, Jinqi, Tang, Zhexiang, Gong, Yu, Ying, Yibiao, Sui, Yang, Lin, Jintian, Huang, Wen, Yuan, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EcoSpa: Efficient Transformer Training with Coupled Sparsity
von: Xiao, Jinqi, et al.
Veröffentlicht: (2025)
von: Xiao, Jinqi, et al.
Veröffentlicht: (2025)
TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
von: Yang, Cheng, et al.
Veröffentlicht: (2025)
von: Yang, Cheng, et al.
Veröffentlicht: (2025)
Action-Guided Attention for Video Action Anticipation
von: Tai, Tsung-Ming, et al.
Veröffentlicht: (2026)
von: Tai, Tsung-Ming, et al.
Veröffentlicht: (2026)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
von: Qian, Zefeng, et al.
Veröffentlicht: (2025)
von: Qian, Zefeng, et al.
Veröffentlicht: (2025)
Quantum Scalar Field Theory Based on the Extended Least Action Principle
von: Yang, Jianhao M.
Veröffentlicht: (2023)
von: Yang, Jianhao M.
Veröffentlicht: (2023)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
Latent Action Control for Reasoning-Guided Unified Image Generation
von: Zhai, Fuxiang, et al.
Veröffentlicht: (2026)
von: Zhai, Fuxiang, et al.
Veröffentlicht: (2026)
VLAMotor: Test-Guided Enhancement of Vision-Language-Action Models via Agent-BasedData Synthesis
von: Liao, Zeqin, et al.
Veröffentlicht: (2026)
von: Liao, Zeqin, et al.
Veröffentlicht: (2026)
COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models
von: Xiao, Jinqi, et al.
Veröffentlicht: (2023)
von: Xiao, Jinqi, et al.
Veröffentlicht: (2023)
Quantum Mechanics Based on an Extended Least Action Principle and Information Metrics of Vacuum Fluctuations
von: Yang, Jianhao M.
Veröffentlicht: (2023)
von: Yang, Jianhao M.
Veröffentlicht: (2023)
Once4All: Skeleton-Guided SMT Solver Fuzzing with LLM-Synthesized Generators
von: Sun, Maolin, et al.
Veröffentlicht: (2025)
von: Sun, Maolin, et al.
Veröffentlicht: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
von: Li, Runze, et al.
Veröffentlicht: (2026)
von: Li, Runze, et al.
Veröffentlicht: (2026)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks
von: Sui, Yang, et al.
Veröffentlicht: (2024)
von: Sui, Yang, et al.
Veröffentlicht: (2024)
Fast Cross-Operator Optimization of Attention Dataflow
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
von: Zhang, Naifu, et al.
Veröffentlicht: (2025)
von: Zhang, Naifu, et al.
Veröffentlicht: (2025)
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
Spin Theory Based on the Extended Least Action Principle and Information Metrics: Quantization, Entanglement, and Bell Test With Time Delay
von: Yang, Jianhao M.
Veröffentlicht: (2024)
von: Yang, Jianhao M.
Veröffentlicht: (2024)
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
von: Guo, Yichen, et al.
Veröffentlicht: (2025)
von: Guo, Yichen, et al.
Veröffentlicht: (2025)
Language Model Guided Interpretable Video Action Reasoning
von: Wang, Ning, et al.
Veröffentlicht: (2024)
von: Wang, Ning, et al.
Veröffentlicht: (2024)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
von: Xu, Changhua, et al.
Veröffentlicht: (2026)
von: Xu, Changhua, et al.
Veröffentlicht: (2026)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
von: Liu, Yudong, et al.
Veröffentlicht: (2026)
von: Liu, Yudong, et al.
Veröffentlicht: (2026)
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models
von: Liu, Haoyun, et al.
Veröffentlicht: (2026)
von: Liu, Haoyun, et al.
Veröffentlicht: (2026)
$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
von: Zhu, Yijie, et al.
Veröffentlicht: (2026)
von: Zhu, Yijie, et al.
Veröffentlicht: (2026)
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
von: Huang, Wenhui, et al.
Veröffentlicht: (2025)
von: Huang, Wenhui, et al.
Veröffentlicht: (2025)
KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
von: Han, Gaoge, et al.
Veröffentlicht: (2026)
von: Han, Gaoge, et al.
Veröffentlicht: (2026)
TAG-HGT: A Scalable and Cost-Effective Framework for Inductive Cold-Start Academic Recommendation
von: Li, Zhexiang
Veröffentlicht: (2025)
von: Li, Zhexiang
Veröffentlicht: (2025)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
Semantically Guided Action Anticipation
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
Self-Guided Action Diffusion
von: Malhotra, Rhea, et al.
Veröffentlicht: (2025)
von: Malhotra, Rhea, et al.
Veröffentlicht: (2025)
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
von: Jia, Xiaosong, et al.
Veröffentlicht: (2026)
von: Jia, Xiaosong, et al.
Veröffentlicht: (2026)
Mask Wearing Fosters Relaxation and Store Engagement in an Offline‐Retail Context
von: Lu Yang, et al.
Veröffentlicht: (2024)
von: Lu Yang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EcoSpa: Efficient Transformer Training with Coupled Sparsity
von: Xiao, Jinqi, et al.
Veröffentlicht: (2025) -
TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
von: Yang, Cheng, et al.
Veröffentlicht: (2025) -
Action-Guided Attention for Video Action Anticipation
von: Tai, Tsung-Ming, et al.
Veröffentlicht: (2026) -
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
von: Yang, Cheng, et al.
Veröffentlicht: (2024) -
Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
von: Qian, Zefeng, et al.
Veröffentlicht: (2025)