V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lin, Xinying, Liu, Xuyang, Wang, Yiyu, Ma, Teng, Ren, Wenqi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
par: Liu, Xuyang, et autres
Publié: (2025)
par: Liu, Xuyang, et autres
Publié: (2025)
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
par: Wang, Yiyu, et autres
Publié: (2025)
par: Wang, Yiyu, et autres
Publié: (2025)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
par: Gao, Shida, et autres
Publié: (2025)
par: Gao, Shida, et autres
Publié: (2025)
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
par: Zhuang, Weijun, et autres
Publié: (2026)
par: Zhuang, Weijun, et autres
Publié: (2026)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
par: Xu, Wenhao, et autres
Publié: (2025)
par: Xu, Wenhao, et autres
Publié: (2025)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
par: Huang, Xiaohu, et autres
Publié: (2024)
par: Huang, Xiaohu, et autres
Publié: (2024)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
par: Cheng, Zixu, et autres
Publié: (2025)
par: Cheng, Zixu, et autres
Publié: (2025)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
par: Wang, Jiankang, et autres
Publié: (2025)
par: Wang, Jiankang, et autres
Publié: (2025)
Temporal Aware Pruning for Efficient Diffusion-based Video Generation
par: Li, Sheng, et autres
Publié: (2026)
par: Li, Sheng, et autres
Publié: (2026)
ShaRP: SHAllow-LayeR Pruning for Efficient Video Large Language Models
par: Xia, Yingjie, et autres
Publié: (2025)
par: Xia, Yingjie, et autres
Publié: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
CAST: Modeling Visual State Transitions for Consistent Video Retrieval
par: Liu, Yanqing, et autres
Publié: (2026)
par: Liu, Yanqing, et autres
Publié: (2026)
Video-Language Alignment via Spatio-Temporal Graph Transformer
par: Zhang, Shi-Xue, et autres
Publié: (2024)
par: Zhang, Shi-Xue, et autres
Publié: (2024)
CAST: Cross-Attentive Spatio-Temporal feature fusion for deepfake detection
par: Thakre, Aryan, et autres
Publié: (2025)
par: Thakre, Aryan, et autres
Publié: (2025)
Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution
par: An, Hongyu, et autres
Publié: (2024)
par: An, Hongyu, et autres
Publié: (2024)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
par: Li, Peiyan, et autres
Publié: (2026)
par: Li, Peiyan, et autres
Publié: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
par: Fu, Honghao, et autres
Publié: (2026)
par: Fu, Honghao, et autres
Publié: (2026)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
par: Meng, Jiahao, et autres
Publié: (2026)
par: Meng, Jiahao, et autres
Publié: (2026)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
par: Sha, Lin, et autres
Publié: (2026)
par: Sha, Lin, et autres
Publié: (2026)
Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video Detection
par: Nguyen, Dat, et autres
Publié: (2025)
par: Nguyen, Dat, et autres
Publié: (2025)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
par: Wu, Jianlong, et autres
Publié: (2025)
par: Wu, Jianlong, et autres
Publié: (2025)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
par: Chen, Junjie, et autres
Publié: (2025)
par: Chen, Junjie, et autres
Publié: (2025)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
par: Cao, Tri, et autres
Publié: (2026)
par: Cao, Tri, et autres
Publié: (2026)
VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
par: Tang, Linfeng, et autres
Publié: (2025)
par: Tang, Linfeng, et autres
Publié: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
par: Sugandhika, Chinthani, et autres
Publié: (2025)
par: Sugandhika, Chinthani, et autres
Publié: (2025)
Spatio-Temporal Correlation Guided Geometric Partitioning for Versatile Video Coding
par: Meng, Xuewei, et autres
Publié: (2026)
par: Meng, Xuewei, et autres
Publié: (2026)
Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency
par: Wang, Yutong, et autres
Publié: (2024)
par: Wang, Yutong, et autres
Publié: (2024)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
par: Li, Jiameng, et autres
Publié: (2026)
par: Li, Jiameng, et autres
Publié: (2026)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
par: Meng, Jiahao, et autres
Publié: (2025)
par: Meng, Jiahao, et autres
Publié: (2025)
FastVID: Dynamic Density Pruning for Fast Video Large Language Models
par: Shen, Leqi, et autres
Publié: (2025)
par: Shen, Leqi, et autres
Publié: (2025)
STGV: Spatio-Temporal Hash Encoding for Gaussian-based Video Representation
par: Lin, Jierun, et autres
Publié: (2026)
par: Lin, Jierun, et autres
Publié: (2026)
VideoMamba: Spatio-Temporal Selective State Space Model
par: Park, Jinyoung, et autres
Publié: (2024)
par: Park, Jinyoung, et autres
Publié: (2024)
Context-Guided Spatio-Temporal Video Grounding
par: Gu, Xin, et autres
Publié: (2024)
par: Gu, Xin, et autres
Publié: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
par: Lan, Xiaohan, et autres
Publié: (2024)
par: Lan, Xiaohan, et autres
Publié: (2024)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
par: Shi, Jiapeng, et autres
Publié: (2026)
par: Shi, Jiapeng, et autres
Publié: (2026)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
par: Liu, Yang, et autres
Publié: (2024)
par: Liu, Yang, et autres
Publié: (2024)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
par: Bagrov, Natan, et autres
Publié: (2025)
par: Bagrov, Natan, et autres
Publié: (2025)
CAST: Cross-Attention in Space and Time for Video Action Recognition
par: Lee, Dongho, et autres
Publié: (2023)
par: Lee, Dongho, et autres
Publié: (2023)
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
par: Jiao, Yingying, et autres
Publié: (2025)
par: Jiao, Yingying, et autres
Publié: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
par: Guo, Yanan, et autres
Publié: (2025)
par: Guo, Yanan, et autres
Publié: (2025)
Documents similaires
-
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
par: Liu, Xuyang, et autres
Publié: (2025) -
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
par: Wang, Yiyu, et autres
Publié: (2025) -
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
par: Gao, Shida, et autres
Publié: (2025) -
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
par: Zhuang, Weijun, et autres
Publié: (2026) -
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
par: Xu, Wenhao, et autres
Publié: (2025)