VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Zihao, Mao, Zhihao, Zhou, Xingyue, Chen, Jiayu, Li, Maoliang, Sun, Xinhao, Zou, Hailong, Zhang, Zhaobo, Liu, Xuanzhe, Cao, Donggang, Mei, Hong, Chen, Xiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
von: Sun, Xinhao, et al.
Veröffentlicht: (2025)
von: Sun, Xinhao, et al.
Veröffentlicht: (2025)
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues
von: Zeng, Tianle, et al.
Veröffentlicht: (2026)
von: Zeng, Tianle, et al.
Veröffentlicht: (2026)
VLN-Zero: Rapid Exploration and Cache-Enabled Neurosymbolic Vision-Language Planning for Zero-Shot Transfer in Robot Navigation
von: Bhatt, Neel P., et al.
Veröffentlicht: (2025)
von: Bhatt, Neel P., et al.
Veröffentlicht: (2025)
ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
IMAC-AgriVLN: Can Agricultural Vision-and-Language Navigation Agents be Aware of Instruction Mistakes?
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2026)
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2026)
T-araVLN: Translator for Agricultural Robotic Agents on Vision-and-Language Navigation
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
Harnessing Input-Adaptive Inference for Efficient VLN
von: Kang, Dongwoo, et al.
Veröffentlicht: (2025)
von: Kang, Dongwoo, et al.
Veröffentlicht: (2025)
AgentVLN: Towards Agentic Vision-and-Language Navigation
von: Xin, Zihao, et al.
Veröffentlicht: (2026)
von: Xin, Zihao, et al.
Veröffentlicht: (2026)
VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation
von: Yu, Bangguo, et al.
Veröffentlicht: (2024)
von: Yu, Bangguo, et al.
Veröffentlicht: (2024)
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
MDE-AgriVLN: Agricultural Vision-and-Language Navigation with Monocular Depth Estimation
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation
von: Ning, Yuwei, et al.
Veröffentlicht: (2026)
von: Ning, Yuwei, et al.
Veröffentlicht: (2026)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
von: Fang, Shaoke, et al.
Veröffentlicht: (2026)
von: Fang, Shaoke, et al.
Veröffentlicht: (2026)
SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
von: Huang, Jingzhi, et al.
Veröffentlicht: (2026)
von: Huang, Jingzhi, et al.
Veröffentlicht: (2026)
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
von: Xin, Zihao, et al.
Veröffentlicht: (2026)
von: Xin, Zihao, et al.
Veröffentlicht: (2026)
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
von: Wen, Jiawen, et al.
Veröffentlicht: (2026)
von: Wen, Jiawen, et al.
Veröffentlicht: (2026)
OmniVLN: Omnidirectional 3D Perception and Token-Efficient LLM Reasoning for Visual-Language Navigation across Air and Ground Platforms
von: Liu, Zhongyuang, et al.
Veröffentlicht: (2026)
von: Liu, Zhongyuang, et al.
Veröffentlicht: (2026)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
von: Yang, Jiahao, et al.
Veröffentlicht: (2026)
von: Yang, Jiahao, et al.
Veröffentlicht: (2026)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN
von: Xia, Ziyi, et al.
Veröffentlicht: (2026)
von: Xia, Ziyi, et al.
Veröffentlicht: (2026)
AIGeN: An Adversarial Approach for Instruction Generation in VLN
von: Rawal, Niyati, et al.
Veröffentlicht: (2024)
von: Rawal, Niyati, et al.
Veröffentlicht: (2024)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
von: Zeng, Shuang, et al.
Veröffentlicht: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
von: Dai, Guangzhao, et al.
Veröffentlicht: (2024)
von: Dai, Guangzhao, et al.
Veröffentlicht: (2024)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
von: Wei, Meng, et al.
Veröffentlicht: (2025)
von: Wei, Meng, et al.
Veröffentlicht: (2025)
AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
von: Loh, Dillon, et al.
Veröffentlicht: (2024)
von: Loh, Dillon, et al.
Veröffentlicht: (2024)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration
von: Cao, Fanpu, et al.
Veröffentlicht: (2025)
von: Cao, Fanpu, et al.
Veröffentlicht: (2025)
OpenVLN: Open-world Aerial Vision-Language Navigation
von: Lin, Peican, et al.
Veröffentlicht: (2025)
von: Lin, Peican, et al.
Veröffentlicht: (2025)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching
von: Zheng, Zihao, et al.
Veröffentlicht: (2026) -
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026) -
HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
von: Zheng, Zihao, et al.
Veröffentlicht: (2026) -
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026) -
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
von: Sun, Xinhao, et al.
Veröffentlicht: (2025)