Small Vision-Language Models are Smart Compressors for Long Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Fei, Junjie, Chen, Jun, Liu, Zechun, Xiong, Yunyang, Zhou, Chong, Wen, Wei, Han, Junlin, Zhuge, Mingchen, Suri, Saksham, Qian, Qi, Liu, Shuming, Wu, Lemeng, Krishnamoorthi, Raghuraman, Chandra, Vikas, Elhoseiny, Mohamed, Zhu, Chenchen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Track Anything
by: Xiong, Yunyang, et al.
Published: (2024)
by: Xiong, Yunyang, et al.
Published: (2024)
EdgeTAM: On-Device Track Anything Model
by: Zhou, Chong, et al.
Published: (2025)
by: Zhou, Chong, et al.
Published: (2025)
PathFusion: Path-consistent Lidar-Camera Deep Feature Fusion
by: Wu, Lemeng, et al.
Published: (2022)
by: Wu, Lemeng, et al.
Published: (2022)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
by: Liu, Shuming, et al.
Published: (2026)
by: Liu, Shuming, et al.
Published: (2026)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
by: Shen, Xiaoqian, et al.
Published: (2024)
by: Shen, Xiaoqian, et al.
Published: (2024)
SqueezeSAM: User friendly mobile interactive segmentation
by: Varadarajan, Balakrishnan, et al.
Published: (2023)
by: Varadarajan, Balakrishnan, et al.
Published: (2023)
Efficient Universal Perception Encoder
by: Zhu, Chenchen, et al.
Published: (2026)
by: Zhu, Chenchen, et al.
Published: (2026)
dTRPO: Trajectory Reduction in Policy Optimization of Diffusion Large Language Models
by: Zhang, Wenxuan, et al.
Published: (2026)
by: Zhang, Wenxuan, et al.
Published: (2026)
MobileMoE: Scaling On-Device Mixture of Experts
by: Chen, Yanbei, et al.
Published: (2026)
by: Chen, Yanbei, et al.
Published: (2026)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
by: Jawahar, Ganesh, et al.
Published: (2023)
by: Jawahar, Ganesh, et al.
Published: (2023)
Agent-as-a-Judge: Evaluate Agents with Agents
by: Zhuge, Mingchen, et al.
Published: (2024)
by: Zhuge, Mingchen, et al.
Published: (2024)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
by: Zhao, Changsheng, et al.
Published: (2025)
by: Zhao, Changsheng, et al.
Published: (2025)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
DepthShrinker: A New Compression Paradigm Towards Boosting Real-Hardware Efficiency of Compact Neural Networks
by: Fu, Yonggan, et al.
Published: (2022)
by: Fu, Yonggan, et al.
Published: (2022)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
by: Fedorov, Igor, et al.
Published: (2024)
by: Fedorov, Igor, et al.
Published: (2024)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
by: Agarwal, Vatsal, et al.
Published: (2026)
by: Agarwal, Vatsal, et al.
Published: (2026)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
by: Ahmed, Mahmoud, et al.
Published: (2024)
by: Ahmed, Mahmoud, et al.
Published: (2024)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
by: Ahmed, Mahmoud, et al.
Published: (2025)
by: Ahmed, Mahmoud, et al.
Published: (2025)
Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models
by: Xiong, Ruibin, et al.
Published: (2025)
by: Xiong, Ruibin, et al.
Published: (2025)
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026)
by: Walmer, Matthew, et al.
Published: (2026)
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)
by: Suri, Saksham, et al.
Published: (2024)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023)
by: Shen, Xiaoqian, et al.
Published: (2023)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
TransCompressor: LLM-Powered Multimodal Data Compression for Smart Transportation
by: Yang, Huanqi, et al.
Published: (2024)
by: Yang, Huanqi, et al.
Published: (2024)
The Compressor-Retriever Architecture for Language Model OS
by: Yang, Yuan, et al.
Published: (2024)
by: Yang, Yuan, et al.
Published: (2024)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Neural Computers
by: Zhuge, Mingchen, et al.
Published: (2026)
by: Zhuge, Mingchen, et al.
Published: (2026)
Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
by: Ji, Ge-Peng, et al.
Published: (2021)
by: Ji, Ge-Peng, et al.
Published: (2021)
M-MiniGPT4: Multilingual VLLM Alignment via Translated Data
by: Han, Seung Hun, et al.
Published: (2026)
by: Han, Seung Hun, et al.
Published: (2026)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
by: Chen, Jun, et al.
Published: (2024)
by: Chen, Jun, et al.
Published: (2024)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024)
by: Wang, Hanyu, et al.
Published: (2024)
Target-Aware Language Modeling via Granular Data Sampling
by: Chang, Ernie, et al.
Published: (2024)
by: Chang, Ernie, et al.
Published: (2024)
XProvence: Zero-Cost Multilingual Context Pruning for Retrieval-Augmented Generation
by: Mohamed, Youssef, et al.
Published: (2026)
by: Mohamed, Youssef, et al.
Published: (2026)
Similar Items
-
Efficient Track Anything
by: Xiong, Yunyang, et al.
Published: (2024) -
EdgeTAM: On-Device Track Anything Model
by: Zhou, Chong, et al.
Published: (2025) -
PathFusion: Path-consistent Lidar-Camera Deep Feature Fusion
by: Wu, Lemeng, et al.
Published: (2022) -
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
by: Liu, Shuming, et al.
Published: (2026) -
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
by: Shen, Xiaoqian, et al.
Published: (2024)