FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Guo, Yanan, Dong, Wenhui, Song, Jun, Zhu, Shiding, Zhang, Xuan, Yang, Hanqing, Wang, Yingbo, Du, Yang, Chen, Xianing, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FILA: Fine-Grained Vision Language Models
par: Zhu, Shiding, et autres
Publié: (2024)
par: Zhu, Shiding, et autres
Publié: (2024)
Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning
par: Ma, Yinchao, et autres
Publié: (2026)
par: Ma, Yinchao, et autres
Publié: (2026)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
par: Liu, Xiangrui, et autres
Publié: (2025)
par: Liu, Xiangrui, et autres
Publié: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
par: Hu, Pengfei, et autres
Publié: (2025)
par: Hu, Pengfei, et autres
Publié: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
par: Fu, Honghao, et autres
Publié: (2026)
par: Fu, Honghao, et autres
Publié: (2026)
L-STEC: Learned Video Compression with Long-term Spatio-Temporal Enhanced Context
par: Zhang, Tiange, et autres
Publié: (2025)
par: Zhang, Tiange, et autres
Publié: (2025)
R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
par: Zhu, Lu, et autres
Publié: (2025)
par: Zhu, Lu, et autres
Publié: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
Towards Long-Form Spatio-Temporal Video Grounding
par: Gu, Xin, et autres
Publié: (2026)
par: Gu, Xin, et autres
Publié: (2026)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
par: Zhou, Xingcheng, et autres
Publié: (2025)
par: Zhou, Xingcheng, et autres
Publié: (2025)
Task-Aware KV Compression For Cost-Effective Long Video Understanding
par: Qin, Minghao, et autres
Publié: (2025)
par: Qin, Minghao, et autres
Publié: (2025)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
par: Zhao, Fufangchen, et autres
Publié: (2025)
par: Zhao, Fufangchen, et autres
Publié: (2025)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
par: Li, Handong, et autres
Publié: (2026)
par: Li, Handong, et autres
Publié: (2026)
Fine-Grained Motion Compression and Selective Temporal Fusion for Neural B-Frame Video Coding
par: Sheng, Xihua, et autres
Publié: (2025)
par: Sheng, Xihua, et autres
Publié: (2025)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
par: Wang, Shaobo, et autres
Publié: (2025)
par: Wang, Shaobo, et autres
Publié: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
par: Wang, Xiao, et autres
Publié: (2026)
par: Wang, Xiao, et autres
Publié: (2026)
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
par: Hu, Yangliu, et autres
Publié: (2025)
par: Hu, Yangliu, et autres
Publié: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
par: Wasim, Syed Talal, et autres
Publié: (2023)
par: Wasim, Syed Talal, et autres
Publié: (2023)
Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features
par: Ji, Shihao, et autres
Publié: (2025)
par: Ji, Shihao, et autres
Publié: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
par: Gupta, Animesh, et autres
Publié: (2025)
par: Gupta, Animesh, et autres
Publié: (2025)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
par: Shen, Xiaoqian, et autres
Publié: (2024)
par: Shen, Xiaoqian, et autres
Publié: (2024)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
par: Cai, Mu, et autres
Publié: (2024)
par: Cai, Mu, et autres
Publié: (2024)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
par: Xu, Wenhao, et autres
Publié: (2025)
par: Xu, Wenhao, et autres
Publié: (2025)
Text-guided Fine-Grained Video Anomaly Understanding
par: Gu, Jihao, et autres
Publié: (2025)
par: Gu, Jihao, et autres
Publié: (2025)
STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
par: Chen, Junyang, et autres
Publié: (2025)
par: Chen, Junyang, et autres
Publié: (2025)
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
par: Li, Xinhao, et autres
Publié: (2025)
par: Li, Xinhao, et autres
Publié: (2025)
MLVU: Benchmarking Multi-task Long Video Understanding
par: Zhou, Junjie, et autres
Publié: (2024)
par: Zhou, Junjie, et autres
Publié: (2024)
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
par: Nie, Ming, et autres
Publié: (2026)
par: Nie, Ming, et autres
Publié: (2026)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
par: Yang, Zhenyu, et autres
Publié: (2025)
par: Yang, Zhenyu, et autres
Publié: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
par: Ding, Yang, et autres
Publié: (2025)
par: Ding, Yang, et autres
Publié: (2025)
SpatioTemporal Difference Network for Video Depth Super-Resolution
par: Wang, Zhengxue, et autres
Publié: (2025)
par: Wang, Zhengxue, et autres
Publié: (2025)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
par: Peng, Yi-Xing, et autres
Publié: (2025)
par: Peng, Yi-Xing, et autres
Publié: (2025)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
par: Kim, Dahun, et autres
Publié: (2025)
par: Kim, Dahun, et autres
Publié: (2025)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
par: Zhang, Runze, et autres
Publié: (2025)
par: Zhang, Runze, et autres
Publié: (2025)
VCA: Video Curious Agent for Long Video Understanding
par: Yang, Zeyuan, et autres
Publié: (2024)
par: Yang, Zeyuan, et autres
Publié: (2024)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
par: Chen, Houlun, et autres
Publié: (2024)
par: Chen, Houlun, et autres
Publié: (2024)
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
par: Kim, Kibum, et autres
Publié: (2026)
par: Kim, Kibum, et autres
Publié: (2026)
EEA: Exploration-Exploitation Agent for Long Video Understanding
par: Yang, Te, et autres
Publié: (2025)
par: Yang, Te, et autres
Publié: (2025)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
par: Wang, Ziyi, et autres
Publié: (2025)
par: Wang, Ziyi, et autres
Publié: (2025)
An Approach to Enriching Surgical Video Datasets for Fine-Grained Spatial-Temporal Understanding of Vision-Language Models
par: Maack, Lennart, et autres
Publié: (2026)
par: Maack, Lennart, et autres
Publié: (2026)
Documents similaires
-
FILA: Fine-Grained Vision Language Models
par: Zhu, Shiding, et autres
Publié: (2024) -
Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning
par: Ma, Yinchao, et autres
Publié: (2026) -
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
par: Liu, Xiangrui, et autres
Publié: (2025) -
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
par: Hu, Pengfei, et autres
Publié: (2025) -
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
par: Fu, Honghao, et autres
Publié: (2026)