Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xiangchen, Zhang, Jinrui, Wang, Teng, Zhang, Haigang, Zheng, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
by: Tang, Yolo Yunlong, et al.
Published: (2023)
by: Tang, Yolo Yunlong, et al.
Published: (2023)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
by: Zhang, Jinrui, et al.
Published: (2024)
by: Zhang, Jinrui, et al.
Published: (2024)
More Pictures Say More: Visual Intersection Network for Open Set Object Detection
by: Dong, Bingcheng, et al.
Published: (2024)
by: Dong, Bingcheng, et al.
Published: (2024)
See More Details: Efficient Image Super-Resolution by Experts Mining
by: Zamfir, Eduard, et al.
Published: (2024)
by: Zamfir, Eduard, et al.
Published: (2024)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
Seeing More with Less: Video Capsule Endoscopy with Multi-Task Learning
by: Werner, Julia, et al.
Published: (2025)
by: Werner, Julia, et al.
Published: (2025)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
by: Geng, Tiantian, et al.
Published: (2024)
by: Geng, Tiantian, et al.
Published: (2024)
\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation
by: Zhu, Weiye, et al.
Published: (2026)
by: Zhu, Weiye, et al.
Published: (2026)
Adapted Center and Scale Prediction: More Stable and More Accurate
by: Wang, Wenhao, et al.
Published: (2020)
by: Wang, Wenhao, et al.
Published: (2020)
The More You See in 2D, the More You Perceive in 3D
by: Han, Xinyang, et al.
Published: (2024)
by: Han, Xinyang, et al.
Published: (2024)
See More, Store Less: Memory-Efficient Resolution for Video Moment Retrieval
by: Jeon, Mingyu, et al.
Published: (2026)
by: Jeon, Mingyu, et al.
Published: (2026)
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning
by: Zeng, Fanhu, et al.
Published: (2026)
by: Zeng, Fanhu, et al.
Published: (2026)
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
by: Lai, Yingxin, et al.
Published: (2026)
by: Lai, Yingxin, et al.
Published: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
by: Li, Siyou, et al.
Published: (2025)
by: Li, Siyou, et al.
Published: (2025)
Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
by: Wang, Shaoguang, et al.
Published: (2025)
by: Wang, Shaoguang, et al.
Published: (2025)
Less is More: Token Context-aware Learning for Object Tracking
by: Xu, Chenlong, et al.
Published: (2025)
by: Xu, Chenlong, et al.
Published: (2025)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
See More, Change Less: Anatomy-Aware Diffusion for Contrast Enhancement
by: Liu, Junqi, et al.
Published: (2025)
by: Liu, Junqi, et al.
Published: (2025)
VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
by: Bian, Jinyue, et al.
Published: (2025)
by: Bian, Jinyue, et al.
Published: (2025)
Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
by: Zhang, Wen, et al.
Published: (2025)
by: Zhang, Wen, et al.
Published: (2025)
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
by: Wang, Huanyu, et al.
Published: (2025)
by: Wang, Huanyu, et al.
Published: (2025)
Less Is More: An Explainable AI Framework for Lightweight Malaria Classification
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
LeMoRe: Learn More Details for Lightweight Semantic Segmentation
by: Abid, Mian Muhammad Naeem, et al.
Published: (2025)
by: Abid, Mian Muhammad Naeem, et al.
Published: (2025)
MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving
by: Zhang, Enming, et al.
Published: (2024)
by: Zhang, Enming, et al.
Published: (2024)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
by: Feng, Mingqian, et al.
Published: (2024)
by: Feng, Mingqian, et al.
Published: (2024)
Let Video Teaches You More: Video-to-Image Knowledge Distillation using DEtection TRansformer for Medical Video Lesion Detection
by: Jiang, Yuncheng, et al.
Published: (2024)
by: Jiang, Yuncheng, et al.
Published: (2024)
Lightweight Adapter Learning for More Generalized Remote Sensing Change Detection
by: Quan, Dou, et al.
Published: (2025)
by: Quan, Dou, et al.
Published: (2025)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
More than the Sum: Panorama-Language Models for Adverse Omni-Scenes
by: Fan, Weijia, et al.
Published: (2026)
by: Fan, Weijia, et al.
Published: (2026)
ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More
by: Zhou, Jiazhou, et al.
Published: (2024)
by: Zhou, Jiazhou, et al.
Published: (2024)
See&Say: Vision Language Guided Safe Zone Detection for Autonomous Package Delivery Drones
by: Ghazanfari, Mahyar, et al.
Published: (2026)
by: Ghazanfari, Mahyar, et al.
Published: (2026)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
Learning More by Seeing Less: Structure First Learning for Efficient, Transferable, and Human-Aligned Vision
by: Li, Tianqin, et al.
Published: (2025)
by: Li, Tianqin, et al.
Published: (2025)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
by: Zhang, Yongchang, et al.
Published: (2026)
by: Zhang, Yongchang, et al.
Published: (2026)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
COLI: A Hierarchical Efficient Compressor for Large Images
by: Wang, Haoran, et al.
Published: (2025)
by: Wang, Haoran, et al.
Published: (2025)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification
by: Zhang, Ke, et al.
Published: (2026)
by: Zhang, Ke, et al.
Published: (2026)
Similar Items
-
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
by: Tang, Yolo Yunlong, et al.
Published: (2023) -
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
by: Zhang, Jinrui, et al.
Published: (2024) -
More Pictures Say More: Visual Intersection Network for Open Set Object Detection
by: Dong, Bingcheng, et al.
Published: (2024) -
See More Details: Efficient Image Super-Resolution by Experts Mining
by: Zamfir, Eduard, et al.
Published: (2024) -
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
by: Wen, Zichen, et al.
Published: (2025)