Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Kibum, Kim, Jiwan, Min, Kyle, Wang, Yueqi, Moon, Jinyoung, McAuley, Julian, Park, Chanyoung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Token-Efficient Item Representation via Images for LLM Recommender Systems
von: Kim, Kibum, et al.
Veröffentlicht: (2025)
von: Kim, Kibum, et al.
Veröffentlicht: (2025)
Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?
von: Kim, Sein, et al.
Veröffentlicht: (2025)
von: Kim, Sein, et al.
Veröffentlicht: (2025)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
von: Kim, Jiwan, et al.
Veröffentlicht: (2026)
von: Kim, Jiwan, et al.
Veröffentlicht: (2026)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
von: Kim, Kibum, et al.
Veröffentlicht: (2024)
Weakly Supervised Video Scene Graph Generation via Natural Language Supervision
von: Kim, Kibum, et al.
Veröffentlicht: (2025)
von: Kim, Kibum, et al.
Veröffentlicht: (2025)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
von: Kim, Kibum, et al.
Veröffentlicht: (2023)
von: Kim, Kibum, et al.
Veröffentlicht: (2023)
Disentangling and Generating Modalities for Recommendation in Missing Modality Scenarios
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
von: Seo, Sangwoo, et al.
Veröffentlicht: (2025)
von: Seo, Sangwoo, et al.
Veröffentlicht: (2025)
Self-Guided Robust Graph Structure Refinement
von: In, Yeonjun, et al.
Veröffentlicht: (2024)
von: In, Yeonjun, et al.
Veröffentlicht: (2024)
Debiased Graph Poisoning Attack via Contrastive Surrogate Objective
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2024)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2024)
Training Robust Graph Neural Networks by Modeling Noise Dependencies
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
FINEST: Stabilizing Recommendations by Rank-Preserving Fine-Tuning
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
ASAP: Attention Sink Anchored Pruning
von: Lee, Jaehyuk, et al.
Veröffentlicht: (2026)
von: Lee, Jaehyuk, et al.
Veröffentlicht: (2026)
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
von: Park, Jinyoung, et al.
Veröffentlicht: (2025)
von: Park, Jinyoung, et al.
Veröffentlicht: (2025)
Extending Input Contexts of Language Models through Training on Segmented Sequences
von: Karypis, Petros, et al.
Veröffentlicht: (2023)
von: Karypis, Petros, et al.
Veröffentlicht: (2023)
CVA: Context-aware Video-text Alignment for Video Temporal Grounding
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
LVCHAT: Facilitating Long Video Comprehension
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
GSPRec: Temporal-Aware Graph Spectral Filtering for Recommendation
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2025)
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2025)
Bridging Conversational and Collaborative Signals for Conversational Recommendation
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2024)
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2024)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025)
von: Kim, Haven, et al.
Veröffentlicht: (2025)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
Semantic Diversity-aware Prototype-based Learning for Unbiased Scene Graph Generation
von: Jeon, Jaehyeong, et al.
Veröffentlicht: (2024)
von: Jeon, Jaehyeong, et al.
Veröffentlicht: (2024)
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
Sink-Aware Pruning for Diffusion Language Models
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Generative AI: A Survey
von: Cao, Yuanjiang, et al.
Veröffentlicht: (2023)
von: Cao, Yuanjiang, et al.
Veröffentlicht: (2023)
VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG
von: Kim, Junkyum, et al.
Veröffentlicht: (2025)
von: Kim, Junkyum, et al.
Veröffentlicht: (2025)
How to Train Data-Efficient LLMs
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
von: Chung, Hyungjin, et al.
Veröffentlicht: (2025)
von: Chung, Hyungjin, et al.
Veröffentlicht: (2025)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
von: Kim, Namho, et al.
Veröffentlicht: (2025)
von: Kim, Namho, et al.
Veröffentlicht: (2025)
Train Once, Deploy Anywhere: Matryoshka Representation Learning for Multimodal Recommendation
von: Wang, Yueqi, et al.
Veröffentlicht: (2024)
von: Wang, Yueqi, et al.
Veröffentlicht: (2024)
Your Causal Self-Attentive Recommender Hosts a Lonely Neighborhood
von: Wang, Yueqi, et al.
Veröffentlicht: (2024)
von: Wang, Yueqi, et al.
Veröffentlicht: (2024)
RA-SGG: Retrieval-Augmented Scene Graph Generation Framework via Multi-Prototype Learning
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2024)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2024)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
von: Long, Phillip, et al.
Veröffentlicht: (2024)
von: Long, Phillip, et al.
Veröffentlicht: (2024)
Three central limit theorems for the unbounded excursion component of a Gaussian field
von: McAuley, Michael
Veröffentlicht: (2024)
von: McAuley, Michael
Veröffentlicht: (2024)
Ähnliche Einträge
-
Token-Efficient Item Representation via Images for LLM Recommender Systems
von: Kim, Kibum, et al.
Veröffentlicht: (2025) -
Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?
von: Kim, Sein, et al.
Veröffentlicht: (2025) -
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
von: Kim, Jiwan, et al.
Veröffentlicht: (2026) -
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025) -
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
von: Kim, Kibum, et al.
Veröffentlicht: (2024)