Sparrow: Text-Anchored Window Attention with Visual-Semantic Glimpsing for Speculative Decoding in Video LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Libo, Zhang, Zhaoning, Hong, Wangyang, Qiao, Peng, Li, Dongsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Speculative Decoding for Autoregressive Video Generation
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024)
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024)
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding
von: Lv, Qitan, et al.
Veröffentlicht: (2026)
von: Lv, Qitan, et al.
Veröffentlicht: (2026)
Few-shot Semantic Encoding and Decoding for Video Surveillance
von: Cheng, Baoping, et al.
Veröffentlicht: (2025)
von: Cheng, Baoping, et al.
Veröffentlicht: (2025)
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering
von: Hagen, Luca, et al.
Veröffentlicht: (2026)
von: Hagen, Luca, et al.
Veröffentlicht: (2026)
Speculative Decoding Reimagined for Multimodal Large Language Models
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
CASCADE: Context-Aware Relaxation for Speculative Image Decoding
von: Yildirim, Selin, et al.
Veröffentlicht: (2026)
von: Yildirim, Selin, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
von: Gao, Yuansheng, et al.
Veröffentlicht: (2026)
von: Gao, Yuansheng, et al.
Veröffentlicht: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking
von: Li, Yunhao, et al.
Veröffentlicht: (2025)
von: Li, Yunhao, et al.
Veröffentlicht: (2025)
Brain2Text Decoding Model Reveals the Neural Mechanisms of Visual Semantic Processing
von: Feng, Feihan, et al.
Veröffentlicht: (2025)
von: Feng, Feihan, et al.
Veröffentlicht: (2025)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
von: Cao, Helin, et al.
Veröffentlicht: (2025)
von: Cao, Helin, et al.
Veröffentlicht: (2025)
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
von: Zhang, Jielu, et al.
Veröffentlicht: (2023)
von: Zhang, Jielu, et al.
Veröffentlicht: (2023)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
SwinECAT: A Transformer-based fundus disease classification model with Shifted Window Attention and Efficient Channel Attention
von: Gu, Peiran, et al.
Veröffentlicht: (2025)
von: Gu, Peiran, et al.
Veröffentlicht: (2025)
UCAgents: Unidirectional Convergence for Visual Evidence Anchored Multi-Agent Medical Decision-Making
von: Feng, Qianhan, et al.
Veröffentlicht: (2025)
von: Feng, Qianhan, et al.
Veröffentlicht: (2025)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2026)
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2026)
Accurate and Fast Compressed Video Captioning
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
von: Chen, Zhifei, et al.
Veröffentlicht: (2025)
von: Chen, Zhifei, et al.
Veröffentlicht: (2025)
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
von: Tan, Haoyue, et al.
Veröffentlicht: (2026)
von: Tan, Haoyue, et al.
Veröffentlicht: (2026)
RISE-Video: Can Video Generators Decode Implicit World Rules?
von: Liu, Mingxin, et al.
Veröffentlicht: (2026)
von: Liu, Mingxin, et al.
Veröffentlicht: (2026)
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
von: Zhu, Jiayin, et al.
Veröffentlicht: (2025)
von: Zhu, Jiayin, et al.
Veröffentlicht: (2025)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
von: Zhou, Qiongyi, et al.
Veröffentlicht: (2024)
von: Zhou, Qiongyi, et al.
Veröffentlicht: (2024)
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
von: Xu, Yinsong, et al.
Veröffentlicht: (2026)
von: Xu, Yinsong, et al.
Veröffentlicht: (2026)
Toward Generalizing Visual Brain Decoding to Unseen Subjects
von: Kong, Xiangtao, et al.
Veröffentlicht: (2024)
von: Kong, Xiangtao, et al.
Veröffentlicht: (2024)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
von: Wu, Hang, et al.
Veröffentlicht: (2026)
von: Wu, Hang, et al.
Veröffentlicht: (2026)
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
von: Zhang, Hongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Hongyu, et al.
Veröffentlicht: (2025)
Ride the Wave: Precision-Allocated Sparse Attention for Smooth Video Generation
von: Zhang, Wentai, et al.
Veröffentlicht: (2026)
von: Zhang, Wentai, et al.
Veröffentlicht: (2026)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
von: Fan, Senran, et al.
Veröffentlicht: (2024)
von: Fan, Senran, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Speculative Decoding for Autoregressive Video Generation
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026) -
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026) -
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025) -
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
von: Jang, Doohyuk, et al.
Veröffentlicht: (2024) -
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding
von: Lv, Qitan, et al.
Veröffentlicht: (2026)