Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Yigui, Wang, Qinglin, Liu, Yang, Liu, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based Customization
von: Liu, Yisu, et al.
Veröffentlicht: (2024)
von: Liu, Yisu, et al.
Veröffentlicht: (2024)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
von: Ma, Jie, et al.
Veröffentlicht: (2023)
von: Ma, Jie, et al.
Veröffentlicht: (2023)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
von: Caselles-Dupré, Hugo, et al.
Veröffentlicht: (2026)
von: Caselles-Dupré, Hugo, et al.
Veröffentlicht: (2026)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
von: Dokme, Atahan, et al.
Veröffentlicht: (2026)
von: Dokme, Atahan, et al.
Veröffentlicht: (2026)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
von: Su, Qile, et al.
Veröffentlicht: (2025)
von: Su, Qile, et al.
Veröffentlicht: (2025)
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
von: Foss, Aaron, et al.
Veröffentlicht: (2025)
von: Foss, Aaron, et al.
Veröffentlicht: (2025)
Dynamic Residual Encoding with Slide-Level Contrastive Learning for End-to-End Whole Slide Image Representation
von: Jin, Jing, et al.
Veröffentlicht: (2025)
von: Jin, Jing, et al.
Veröffentlicht: (2025)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
von: Han, Yudong, et al.
Veröffentlicht: (2024)
von: Han, Yudong, et al.
Veröffentlicht: (2024)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
von: Safdar, Aon, et al.
Veröffentlicht: (2025)
von: Safdar, Aon, et al.
Veröffentlicht: (2025)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
von: Teo, Charlton
Veröffentlicht: (2025)
von: Teo, Charlton
Veröffentlicht: (2025)
MFTF: Mask-free Training-free Object Level Layout Control Diffusion Model
von: Yang, Shan
Veröffentlicht: (2024)
von: Yang, Shan
Veröffentlicht: (2024)
PhysVid: Physics Aware Local Conditioning for Generative Video Models
von: Pathak, Saurabh, et al.
Veröffentlicht: (2026)
von: Pathak, Saurabh, et al.
Veröffentlicht: (2026)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
von: Sanders, Kate, et al.
Veröffentlicht: (2024)
von: Sanders, Kate, et al.
Veröffentlicht: (2024)
A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
von: Lucas, Andrea Filiberto, et al.
Veröffentlicht: (2026)
von: Lucas, Andrea Filiberto, et al.
Veröffentlicht: (2026)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
Detecting Inpainted Video with Frequency Domain Insights
von: Tang, Quanhui, et al.
Veröffentlicht: (2024)
von: Tang, Quanhui, et al.
Veröffentlicht: (2024)
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
von: Wang, Wentao, et al.
Veröffentlicht: (2025)
von: Wang, Wentao, et al.
Veröffentlicht: (2025)
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
von: Huang, Yuhang, et al.
Veröffentlicht: (2025)
von: Huang, Yuhang, et al.
Veröffentlicht: (2025)
A Simple Baseline for Streaming Video Understanding
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
EDSNet: Efficient-DSNet for Video Summarization
von: Prasad, Ashish, et al.
Veröffentlicht: (2024)
von: Prasad, Ashish, et al.
Veröffentlicht: (2024)
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
von: Hansen-Estruch, Philippe, et al.
Veröffentlicht: (2025)
von: Hansen-Estruch, Philippe, et al.
Veröffentlicht: (2025)
Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective
von: Luo, Wang, et al.
Veröffentlicht: (2025)
von: Luo, Wang, et al.
Veröffentlicht: (2025)
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
von: Atighehchian, Parmida, et al.
Veröffentlicht: (2026)
von: Atighehchian, Parmida, et al.
Veröffentlicht: (2026)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
von: Naeen, Mohammad Ali Etemadi, et al.
Veröffentlicht: (2025)
von: Naeen, Mohammad Ali Etemadi, et al.
Veröffentlicht: (2025)
S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction
von: Adiban, Mohammad, et al.
Veröffentlicht: (2023)
von: Adiban, Mohammad, et al.
Veröffentlicht: (2023)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Addressing Issues with Working Memory in Video Object Segmentation
von: Bromley, Clayton, et al.
Veröffentlicht: (2024)
von: Bromley, Clayton, et al.
Veröffentlicht: (2024)
Hardware-Aware YOLO Compression for Low-Power Edge AI on STM32U5 for Weeds Detection in Digital Agriculture
von: Kouzinopoulos, Charalampos S., et al.
Veröffentlicht: (2025)
von: Kouzinopoulos, Charalampos S., et al.
Veröffentlicht: (2025)
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
von: Agarwal, Rachit, et al.
Veröffentlicht: (2026)
von: Agarwal, Rachit, et al.
Veröffentlicht: (2026)
SITUATE -- Synthetic Object Counting Dataset for VLM training
von: Peinl, René, et al.
Veröffentlicht: (2026)
von: Peinl, René, et al.
Veröffentlicht: (2026)
Image Segmentation and Classification of E-waste for Training Robots for Waste Segregation
von: Tripathi, Prakriti
Veröffentlicht: (2025)
von: Tripathi, Prakriti
Veröffentlicht: (2025)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
von: Holm, Felix, et al.
Veröffentlicht: (2025)
von: Holm, Felix, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
von: Chen, Yuhao, et al.
Veröffentlicht: (2026) -
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026) -
Sora as a World Model? A Complete Survey on Text-to-Video Generation
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024) -
Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based Customization
von: Liu, Yisu, et al.
Veröffentlicht: (2024) -
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025)