V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Xinying, Liu, Xuyang, Wang, Yiyu, Ma, Teng, Ren, Wenqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
von: Wang, Yiyu, et al.
Veröffentlicht: (2025)
von: Wang, Yiyu, et al.
Veröffentlicht: (2025)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
von: Gao, Shida, et al.
Veröffentlicht: (2025)
von: Gao, Shida, et al.
Veröffentlicht: (2025)
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
von: Zhuang, Weijun, et al.
Veröffentlicht: (2026)
von: Zhuang, Weijun, et al.
Veröffentlicht: (2026)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
Temporal Aware Pruning for Efficient Diffusion-based Video Generation
von: Li, Sheng, et al.
Veröffentlicht: (2026)
von: Li, Sheng, et al.
Veröffentlicht: (2026)
ShaRP: SHAllow-LayeR Pruning for Efficient Video Large Language Models
von: Xia, Yingjie, et al.
Veröffentlicht: (2025)
von: Xia, Yingjie, et al.
Veröffentlicht: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
von: Kumar, Yogesh
Veröffentlicht: (2025)
von: Kumar, Yogesh
Veröffentlicht: (2025)
CAST: Modeling Visual State Transitions for Consistent Video Retrieval
von: Liu, Yanqing, et al.
Veröffentlicht: (2026)
von: Liu, Yanqing, et al.
Veröffentlicht: (2026)
Video-Language Alignment via Spatio-Temporal Graph Transformer
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
CAST: Cross-Attentive Spatio-Temporal feature fusion for deepfake detection
von: Thakre, Aryan, et al.
Veröffentlicht: (2025)
von: Thakre, Aryan, et al.
Veröffentlicht: (2025)
Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution
von: An, Hongyu, et al.
Veröffentlicht: (2024)
von: An, Hongyu, et al.
Veröffentlicht: (2024)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
von: Li, Peiyan, et al.
Veröffentlicht: (2026)
von: Li, Peiyan, et al.
Veröffentlicht: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
von: Sha, Lin, et al.
Veröffentlicht: (2026)
von: Sha, Lin, et al.
Veröffentlicht: (2026)
Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection
von: Nguyen, Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Dat, et al.
Veröffentlicht: (2025)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
Spatio-Temporal Correlation Guided Geometric Partitioning for Versatile Video Coding
von: Meng, Xuewei, et al.
Veröffentlicht: (2026)
von: Meng, Xuewei, et al.
Veröffentlicht: (2026)
Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
von: Li, Jiameng, et al.
Veröffentlicht: (2026)
von: Li, Jiameng, et al.
Veröffentlicht: (2026)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
FastVID: Dynamic Density Pruning for Fast Video Large Language Models
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
STGV: Spatio-Temporal Hash Encoding for Gaussian-based Video Representation
von: Lin, Jierun, et al.
Veröffentlicht: (2026)
von: Lin, Jierun, et al.
Veröffentlicht: (2026)
VideoMamba: Spatio-Temporal Selective State Space Model
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
CAST: Cross-Attention in Space and Time for Video Action Recognition
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
von: Jiao, Yingying, et al.
Veröffentlicht: (2025)
von: Jiao, Yingying, et al.
Veröffentlicht: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025) -
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
von: Wang, Yiyu, et al.
Veröffentlicht: (2025) -
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
von: Gao, Shida, et al.
Veröffentlicht: (2025) -
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
von: Zhuang, Weijun, et al.
Veröffentlicht: (2026) -
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)