StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Tongqing, Wu, Hang, Wang, Jiasen, Li, Xiaotao, Fang, Lu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SuperSuit: An Isomorphic Bimodal Interface for Scalable Mobile Manipulation
von: Chen, Tongqing, et al.
Veröffentlicht: (2026)
von: Chen, Tongqing, et al.
Veröffentlicht: (2026)
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
MolmoAct: Action Reasoning Models that can Reason in Space
von: Lee, Jason, et al.
Veröffentlicht: (2025)
von: Lee, Jason, et al.
Veröffentlicht: (2025)
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving
von: Huang, Yuzhou, et al.
Veröffentlicht: (2026)
von: Huang, Yuzhou, et al.
Veröffentlicht: (2026)
ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
von: Zhang, Zhuoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhuoyang, et al.
Veröffentlicht: (2026)
BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning
von: Wang, Ruiheng, et al.
Veröffentlicht: (2026)
von: Wang, Ruiheng, et al.
Veröffentlicht: (2026)
VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts
von: Jiang, Yuhua, et al.
Veröffentlicht: (2026)
von: Jiang, Yuhua, et al.
Veröffentlicht: (2026)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
von: Miao, Cui, et al.
Veröffentlicht: (2025)
von: Miao, Cui, et al.
Veröffentlicht: (2025)
CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
von: Ma, Chenyang, et al.
Veröffentlicht: (2026)
von: Ma, Chenyang, et al.
Veröffentlicht: (2026)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
von: Fan, Xianzhe, et al.
Veröffentlicht: (2026)
von: Fan, Xianzhe, et al.
Veröffentlicht: (2026)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
ReMem-VLA: Empowering Vision-Language-Action Model with Memory via Dual-Level Recurrent Queries
von: Li, Hang, et al.
Veröffentlicht: (2026)
von: Li, Hang, et al.
Veröffentlicht: (2026)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
von: Fang, Zhen, et al.
Veröffentlicht: (2025)
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
von: Bu, Qingwen, et al.
Veröffentlicht: (2025)
von: Bu, Qingwen, et al.
Veröffentlicht: (2025)
Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
von: Jabbour, Jason, et al.
Veröffentlicht: (2025)
von: Jabbour, Jason, et al.
Veröffentlicht: (2025)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
ReTac-ACT: A State-Gated Vision-Tactile Fusion Transformer for Precision Assembly
von: Ruan, Minchi, et al.
Veröffentlicht: (2026)
von: Ruan, Minchi, et al.
Veröffentlicht: (2026)
Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO
von: Chen, Tianyang, et al.
Veröffentlicht: (2026)
von: Chen, Tianyang, et al.
Veröffentlicht: (2026)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
von: Luo, Yuankai, et al.
Veröffentlicht: (2026)
von: Luo, Yuankai, et al.
Veröffentlicht: (2026)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
von: Tur, Yalcin, et al.
Veröffentlicht: (2026)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training
von: Huang, Suning, et al.
Veröffentlicht: (2026)
von: Huang, Suning, et al.
Veröffentlicht: (2026)
FrameSkip: Learning from Fewer but More Informative Frames in VLA Training
von: Yu, Bin, et al.
Veröffentlicht: (2026)
von: Yu, Bin, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Robots via Retrieve-Reason-Act
von: Temiraliev, Izat, et al.
Veröffentlicht: (2026)
von: Temiraliev, Izat, et al.
Veröffentlicht: (2026)
CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
von: Fang, Shiyu, et al.
Veröffentlicht: (2025)
von: Fang, Shiyu, et al.
Veröffentlicht: (2025)
TeleGate: Whole-Body Humanoid Teleoperation via Gated Expert Selection with Motion Prior
von: Li, Jie, et al.
Veröffentlicht: (2026)
von: Li, Jie, et al.
Veröffentlicht: (2026)
Agile-VLA: Few-Shot Industrial Pose Rectification via Implicit Affordance Anchoring
von: Yan, Teng, et al.
Veröffentlicht: (2026)
von: Yan, Teng, et al.
Veröffentlicht: (2026)
DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
von: Shang, Shuyao, et al.
Veröffentlicht: (2026)
von: Shang, Shuyao, et al.
Veröffentlicht: (2026)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
MolmoAct2: Action Reasoning Models for Real-world Deployment
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
von: Fang, Haoquan, et al.
Veröffentlicht: (2026)
UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
TacMamba: A Tactile History Compression Adapter Bridging Fast Reflexes and Slow VLA Reasoning
von: Wang, Zhenan, et al.
Veröffentlicht: (2026)
von: Wang, Zhenan, et al.
Veröffentlicht: (2026)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
von: Chen, Jiahong, et al.
Veröffentlicht: (2025)
von: Chen, Jiahong, et al.
Veröffentlicht: (2025)
Long-Horizon Manipulation via Trace-Conditioned VLA Planning
von: Liu, Isabella, et al.
Veröffentlicht: (2026)
von: Liu, Isabella, et al.
Veröffentlicht: (2026)
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
von: Du, Zhaohui, et al.
Veröffentlicht: (2026)
von: Du, Zhaohui, et al.
Veröffentlicht: (2026)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SuperSuit: An Isomorphic Bimodal Interface for Scalable Mobile Manipulation
von: Chen, Tongqing, et al.
Veröffentlicht: (2026) -
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
von: Guo, Wenkai, et al.
Veröffentlicht: (2025) -
MolmoAct: Action Reasoning Models that can Reason in Space
von: Lee, Jason, et al.
Veröffentlicht: (2025) -
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving
von: Huang, Yuzhou, et al.
Veröffentlicht: (2026) -
ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
von: Zhang, Zhuoyang, et al.
Veröffentlicht: (2026)