Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yudong, Li, Yuan, Tang, Zijia, Zheng, Yuxi, Lin, Yueqian, Wang, Qinsi, Li, Yi, Liu, Shuangjun, Zhang, Shuai, Jing, Taotao, Gao, Dashan, Bi, Ning, Sun, Jingwei, Chen, Yiran, Li, Hai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
Bridging the Perception Gap: A Lightweight Coarse-to-Fine Architecture for Edge Audio Systems
von: Zhang, Hengfan, et al.
Veröffentlicht: (2026)
von: Zhang, Hengfan, et al.
Veröffentlicht: (2026)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion
von: Lin, Yueqian, et al.
Veröffentlicht: (2023)
von: Lin, Yueqian, et al.
Veröffentlicht: (2023)
FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching
von: Fu, Yuzhe, et al.
Veröffentlicht: (2026)
von: Fu, Yuzhe, et al.
Veröffentlicht: (2026)
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
von: Wei, Chiyue, et al.
Veröffentlicht: (2025)
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
von: Wang, Zhongxi, et al.
Veröffentlicht: (2026)
von: Wang, Zhongxi, et al.
Veröffentlicht: (2026)
Learning Generalizable Policy for Obstacle-Aware Autonomous Drone Racing
von: Liu, Yueqian
Veröffentlicht: (2024)
von: Liu, Yueqian
Veröffentlicht: (2024)
Federated Unsupervised Visual Representation Learning via Exploiting General Content and Personal Style
von: Yang, Yuewei, et al.
Veröffentlicht: (2022)
von: Yang, Yuewei, et al.
Veröffentlicht: (2022)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
Bridging the Gap Between Sparsity and Redundancy: A Dual-Decoding Framework with Global Context for Map Inference
von: Shen, Yudong, et al.
Veröffentlicht: (2025)
von: Shen, Yudong, et al.
Veröffentlicht: (2025)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
von: Lin, Yihan, et al.
Veröffentlicht: (2026)
CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
von: Wang, Qinsi, et al.
Veröffentlicht: (2024)
von: Wang, Qinsi, et al.
Veröffentlicht: (2024)
Modeling the Electrical Stabilities of Organic Light‐Emitting Diodes
von: Dashan Qin, et al.
Veröffentlicht: (2024)
von: Dashan Qin, et al.
Veröffentlicht: (2024)
Geometric Stratification for Singular Configurations of the P3P Problem via Local Dual Space
von: Sun, Xueying, et al.
Veröffentlicht: (2026)
von: Sun, Xueying, et al.
Veröffentlicht: (2026)
GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
von: Navardi, Mozhgan, et al.
Veröffentlicht: (2025)
von: Navardi, Mozhgan, et al.
Veröffentlicht: (2025)
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
von: Yue, Tongtian, et al.
Veröffentlicht: (2025)
von: Yue, Tongtian, et al.
Veröffentlicht: (2025)
Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
von: Li, Shuai, et al.
Veröffentlicht: (2025)
von: Li, Shuai, et al.
Veröffentlicht: (2025)
Controllable generation and spatial phase-modulation of vortex beams in cascade-type atomic ensembles
von: Li, Yueqian, et al.
Veröffentlicht: (2024)
von: Li, Yueqian, et al.
Veröffentlicht: (2024)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
von: Li, Zuolei, et al.
Veröffentlicht: (2025)
von: Li, Zuolei, et al.
Veröffentlicht: (2025)
KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
von: Ye, Hancheng, et al.
Veröffentlicht: (2025)
von: Ye, Hancheng, et al.
Veröffentlicht: (2025)
PrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference Training
von: Cheng, Yuhan, et al.
Veröffentlicht: (2026)
von: Cheng, Yuhan, et al.
Veröffentlicht: (2026)
Dual Ion Regulated Eutectogels with High Elasticity and Adhesive Strength for Accurate Strain Sensors
von: Yaozhou Sun, et al.
Veröffentlicht: (2024)
von: Yaozhou Sun, et al.
Veröffentlicht: (2024)
Beyond Static Question Banks: Dynamic Knowledge Expansion via LLM-Automated Graph Construction and Adaptive Generation
von: Wang, Yingquan, et al.
Veröffentlicht: (2026)
von: Wang, Yingquan, et al.
Veröffentlicht: (2026)
AUEditNet: Dual-Branch Facial Action Unit Intensity Manipulation with Implicit Disentanglement
von: Jin, Shiwei, et al.
Veröffentlicht: (2024)
von: Jin, Shiwei, et al.
Veröffentlicht: (2024)
Towards Black-Box Membership Inference Attack for Diffusion Models
von: Li, Jingwei, et al.
Veröffentlicht: (2024)
von: Li, Jingwei, et al.
Veröffentlicht: (2024)
Latent Action Reparameterization for Efficient Agent Inference
von: Huang, Wenhao, et al.
Veröffentlicht: (2026)
von: Huang, Wenhao, et al.
Veröffentlicht: (2026)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
Featured Cover: Cover Image, Volume 90, Issue 11
von: Huisen Wang, et al.
Veröffentlicht: (2025)
von: Huisen Wang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
von: Liu, Yudong, et al.
Veröffentlicht: (2025) -
LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
von: Liu, Yudong, et al.
Veröffentlicht: (2025) -
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
von: Lin, Yueqian, et al.
Veröffentlicht: (2025) -
CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models
von: Wang, Qinsi, et al.
Veröffentlicht: (2025) -
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
von: Wang, Qinsi, et al.
Veröffentlicht: (2025)