FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Tianyu, Liu, Tengxuan, Han, Qinghao, Dai, Guohao, Yan, Shengen, Yang, Huazhong, Ning, Xuefei, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
The MSR-Video to Text Dataset with Clean Annotations
by: Chen, Haoran, et al.
Published: (2021)
by: Chen, Haoran, et al.
Published: (2021)
Context-Dependent Affordance Computation in Vision-Language Models
by: Farzulla, Murad
Published: (2026)
by: Farzulla, Murad
Published: (2026)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
by: Huo, Dongjie, et al.
Published: (2026)
by: Huo, Dongjie, et al.
Published: (2026)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
by: Zhang, Junwen, et al.
Published: (2025)
by: Zhang, Junwen, et al.
Published: (2025)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
A Survey on Vision-Language-Action Models for Embodied AI
by: Ma, Yueen, et al.
Published: (2024)
by: Ma, Yueen, et al.
Published: (2024)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
by: Qi, Dekang, et al.
Published: (2026)
by: Qi, Dekang, et al.
Published: (2026)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
by: Yanambakkam, Hemanth Teja, et al.
Published: (2025)
by: Yanambakkam, Hemanth Teja, et al.
Published: (2025)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
by: Gkountouras, John, et al.
Published: (2025)
by: Gkountouras, John, et al.
Published: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
by: Allakhverdov, Eduard, et al.
Published: (2025)
by: Allakhverdov, Eduard, et al.
Published: (2025)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
by: Yang, Baoyao, et al.
Published: (2025)
by: Yang, Baoyao, et al.
Published: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
by: Sáez, Arnau Igualde, et al.
Published: (2025)
by: Sáez, Arnau Igualde, et al.
Published: (2025)
Does CLIP perceive art the same way we do?
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
by: Koh, Hyunseo, et al.
Published: (2026)
by: Koh, Hyunseo, et al.
Published: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
by: Chen, Kewei, et al.
Published: (2025)
by: Chen, Kewei, et al.
Published: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
by: Zhang, Junbin, et al.
Published: (2022)
by: Zhang, Junbin, et al.
Published: (2022)
Tricks and Plug-ins for Gradient Boosting in Image Classification
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
by: Papyan, Narek, et al.
Published: (2024)
by: Papyan, Narek, et al.
Published: (2024)
ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
by: Liu, Xueyi, et al.
Published: (2025)
by: Liu, Xueyi, et al.
Published: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
by: Zeng, Zhitao, et al.
Published: (2026)
by: Zeng, Zhitao, et al.
Published: (2026)
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
by: Malikussaid, et al.
Published: (2026)
by: Malikussaid, et al.
Published: (2026)
RACAS: Controlling Diverse Robots With a Single Agentic System
by: Ashley, Dylan R., et al.
Published: (2026)
by: Ashley, Dylan R., et al.
Published: (2026)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
by: Cui, Hejie, et al.
Published: (2024)
by: Cui, Hejie, et al.
Published: (2024)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
by: Ouyang, Rongxin, et al.
Published: (2024)
by: Ouyang, Rongxin, et al.
Published: (2024)
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
by: Montello, Fabio, et al.
Published: (2025)
by: Montello, Fabio, et al.
Published: (2025)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
Multimodal Structure-Aware Quantum Data Processing
by: Hawashin, Hala, et al.
Published: (2024)
by: Hawashin, Hala, et al.
Published: (2024)
Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
by: Moore, Alexander, et al.
Published: (2025)
by: Moore, Alexander, et al.
Published: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
by: Patel, Urjitkumar, et al.
Published: (2025)
by: Patel, Urjitkumar, et al.
Published: (2025)
ClustViT: Clustering-based Token Merging for Semantic Segmentation
by: Montello, Fabio, et al.
Published: (2025)
by: Montello, Fabio, et al.
Published: (2025)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025)
by: Zinnen, Mathias, et al.
Published: (2025)
Revisiting Energy-Based Model for Out-of-Distribution Detection
by: Wu, Yifan, et al.
Published: (2024)
by: Wu, Yifan, et al.
Published: (2024)
Similar Items
-
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025) -
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025) -
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025) -
The MSR-Video to Text Dataset with Clean Annotations
by: Chen, Haoran, et al.
Published: (2021) -
Context-Dependent Affordance Computation in Vision-Language Models
by: Farzulla, Murad
Published: (2026)