HierSum: A Global and Local Attention Mechanism for Video Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Beedu, Apoorva, Essa, Irfan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024)
by: Samel, Karan, et al.
Published: (2024)
On the Efficacy of Text-Based Input Modalities for Action Anticipation
by: Beedu, Apoorva, et al.
Published: (2024)
by: Beedu, Apoorva, et al.
Published: (2024)
Mamba Fusion: Learning Actions Through Questioning
by: Dong, Zhikang, et al.
Published: (2024)
by: Dong, Zhikang, et al.
Published: (2024)
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
by: Haresamudram, Harish, et al.
Published: (2024)
by: Haresamudram, Harish, et al.
Published: (2024)
Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
by: Colagrande, Alex, et al.
Published: (2025)
by: Colagrande, Alex, et al.
Published: (2025)
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
by: Yan, Haolong, et al.
Published: (2025)
by: Yan, Haolong, et al.
Published: (2025)
LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
by: Patel, Alkesh, et al.
Published: (2026)
by: Patel, Alkesh, et al.
Published: (2026)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026)
by: Kim, Sumin, et al.
Published: (2026)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
by: Abbasi, Mehryar, et al.
Published: (2024)
by: Abbasi, Mehryar, et al.
Published: (2024)
VideoNSA: Native Sparse Attention Scales Video Understanding
by: Song, Enxin, et al.
Published: (2025)
by: Song, Enxin, et al.
Published: (2025)
Grad Queue : A probabilistic framework to reinforce sparse gradients
by: Hasib, Irfan Mohammad Al
Published: (2024)
by: Hasib, Irfan Mohammad Al
Published: (2024)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
by: Miyato, Takeru, et al.
Published: (2023)
by: Miyato, Takeru, et al.
Published: (2023)
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
by: Breuer, Adam, et al.
Published: (2025)
by: Breuer, Adam, et al.
Published: (2025)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
TempoControl: Temporal Attention Guidance for Text-to-Video Models
by: Schiber, Shira, et al.
Published: (2025)
by: Schiber, Shira, et al.
Published: (2025)
MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
by: Chen, Chieh-Yun, et al.
Published: (2025)
by: Chen, Chieh-Yun, et al.
Published: (2025)
DFA-CON: A Contrastive Learning Approach for Detecting Copyright Infringement in DeepFake Art
by: Wahab, Haroon, et al.
Published: (2025)
by: Wahab, Haroon, et al.
Published: (2025)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
by: Zou, Shihao, et al.
Published: (2025)
by: Zou, Shihao, et al.
Published: (2025)
EdgeVidSum: Real-Time Personalized Video Summarization at the Edge
by: Mujtaba, Ghulam, et al.
Published: (2025)
by: Mujtaba, Ghulam, et al.
Published: (2025)
Locally Testing Model Detections for Semantic Global Concepts
by: Motzkus, Franz, et al.
Published: (2024)
by: Motzkus, Franz, et al.
Published: (2024)
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
by: Seo, Junwon, et al.
Published: (2026)
by: Seo, Junwon, et al.
Published: (2026)
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
by: Gu, Youping, et al.
Published: (2025)
by: Gu, Youping, et al.
Published: (2025)
Generative Dataset Distillation: Balancing Global Structure and Local Details
by: Li, Longzhen, et al.
Published: (2024)
by: Li, Longzhen, et al.
Published: (2024)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
by: Marmon, Andrew, et al.
Published: (2024)
by: Marmon, Andrew, et al.
Published: (2024)
IBiT: Utilizing Inductive Biases to Create a More Data Efficient Attention Mechanism
by: Giri, Adithya
Published: (2025)
by: Giri, Adithya
Published: (2025)
Aggregating Local Saliency Maps for Semi-Global Explainable Image Classification
by: Hinns, James, et al.
Published: (2025)
by: Hinns, James, et al.
Published: (2025)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
Knee-xRAI: An Explainable AI Framework for Automatic Kellgren-Lawrence Grading of Knee Osteoarthritis
by: Irfan, Azmul A., et al.
Published: (2026)
by: Irfan, Azmul A., et al.
Published: (2026)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)
by: Jun, Youngjun, et al.
Published: (2026)
MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
by: Liu, YiTong, et al.
Published: (2025)
by: Liu, YiTong, et al.
Published: (2025)
NULLBUS: Multimodal Mixed-Supervision for Breast Ultrasound Segmentation via Nullable Global-Local Prompts
by: Mallina, Raja, et al.
Published: (2025)
by: Mallina, Raja, et al.
Published: (2025)
An Interpretable Implicit-Based Approach for Modeling Local Spatial Effects: A Case Study of Global Gross Primary Productivity
by: Du, Siqi, et al.
Published: (2025)
by: Du, Siqi, et al.
Published: (2025)
GLC++: Source-Free Universal Domain Adaptation through Global-Local Clustering and Contrastive Affinity Learning
by: Qu, Sanqing, et al.
Published: (2024)
by: Qu, Sanqing, et al.
Published: (2024)
FedDistill: Global Model Distillation for Local Model De-Biasing in Non-IID Federated Learning
by: Song, Changlin, et al.
Published: (2024)
by: Song, Changlin, et al.
Published: (2024)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
by: Ntrougkas, Mariano V., et al.
Published: (2024)
by: Ntrougkas, Mariano V., et al.
Published: (2024)
CUE-Net: Violence Detection Video Analytics with Spatial Cropping, Enhanced UniformerV2 and Modified Efficient Additive Attention
by: Senadeera, Damith Chamalke, et al.
Published: (2024)
by: Senadeera, Damith Chamalke, et al.
Published: (2024)
Attention in Diffusion Model: A Survey
by: Hua, Litao, et al.
Published: (2025)
by: Hua, Litao, et al.
Published: (2025)
Similar Items
-
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024) -
On the Efficacy of Text-Based Input Modalities for Action Anticipation
by: Beedu, Apoorva, et al.
Published: (2024) -
Mamba Fusion: Learning Actions Through Questioning
by: Dong, Zhikang, et al.
Published: (2024) -
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
by: Haresamudram, Harish, et al.
Published: (2024) -
Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
by: Colagrande, Alex, et al.
Published: (2025)