Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Tariq, Omer, Raza, Syed Muhammad, Son, Jeongbae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CSTA: CNN-based Spatiotemporal Attention for Video Summarization
by: Son, Jaewon, et al.
Published: (2024)
by: Son, Jaewon, et al.
Published: (2024)
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Scaling Up Video Summarization Pretraining with Large Language Models
by: Argaw, Dawit Mureja, et al.
Published: (2024)
by: Argaw, Dawit Mureja, et al.
Published: (2024)
EdgeDAM: Real-time Object Tracking for Mobile Devices
by: Raza, Syed Muhammad, et al.
Published: (2026)
by: Raza, Syed Muhammad, et al.
Published: (2026)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
by: Ali, Muhammad, et al.
Published: (2024)
by: Ali, Muhammad, et al.
Published: (2024)
Language-Guided Graph Representation Learning for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
by: Nazir, Maham, et al.
Published: (2026)
by: Nazir, Maham, et al.
Published: (2026)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
by: Sarowar, Md Selim, et al.
Published: (2026)
by: Sarowar, Md Selim, et al.
Published: (2026)
Faithful Counterfactual Visual Explanations (FCVE)
by: Khan, Bismillah, et al.
Published: (2025)
by: Khan, Bismillah, et al.
Published: (2025)
VideoSAGE: Video Summarization with Graph Representation Learning
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
Small Object Detection with YOLO: A Performance Analysis Across Model Versions and Hardware
by: Tariq, Muhammad Fasih, et al.
Published: (2025)
by: Tariq, Muhammad Fasih, et al.
Published: (2025)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
by: Wu, Yuanli, et al.
Published: (2025)
by: Wu, Yuanli, et al.
Published: (2025)
Comparing Learning Paradigms for Egocentric Video Summarization
by: Wen, Daniel
Published: (2025)
by: Wen, Daniel
Published: (2025)
Video Summarization Techniques: A Comprehensive Review
by: Alaa, Toqa, et al.
Published: (2024)
by: Alaa, Toqa, et al.
Published: (2024)
Enhanced Multimodal Content Moderation of Children's Videos using Audiovisual Fusion
by: Ahmed, Syed Hammad, et al.
Published: (2024)
by: Ahmed, Syed Hammad, et al.
Published: (2024)
GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval
by: Jeon, Mingyu, et al.
Published: (2026)
by: Jeon, Mingyu, et al.
Published: (2026)
Uncertainty-Aware Concept and Motion Segmentation for Semi-Supervised Angiography Videos
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Towards Minimal Focal Stack in Shape from Focus
by: Ashfaq, Khurram, et al.
Published: (2026)
by: Ashfaq, Khurram, et al.
Published: (2026)
Robust Shape from Focus via Multiscale Directional Dilated Laplacian and Recurrent Network
by: Ashfaq, Khurram, et al.
Published: (2025)
by: Ashfaq, Khurram, et al.
Published: (2025)
LFRA-Net: A Lightweight Focal and Region-Aware Attention Network for Retinal Vessel Segmentatio
by: Mehmood, Mehwish, et al.
Published: (2025)
by: Mehmood, Mehwish, et al.
Published: (2025)
Leveraging counterfactual concepts for debugging and improving CNN model performance
by: Tariq, Syed Ali, et al.
Published: (2025)
by: Tariq, Syed Ali, et al.
Published: (2025)
Video Summarization using Denoising Diffusion Probabilistic Model
by: Shang, Zirui, et al.
Published: (2024)
by: Shang, Zirui, et al.
Published: (2024)
Spiking Variational Graph Representation Inference for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Stride-Net: Fairness-Aware Disentangled Representation Learning for Chest X-Ray Diagnosis
by: Rashid, Darakshan, et al.
Published: (2026)
by: Rashid, Darakshan, et al.
Published: (2026)
Uncertainty Visualization via Low-Dimensional Posterior Projections
by: Yair, Omer, et al.
Published: (2023)
by: Yair, Omer, et al.
Published: (2023)
CATVis: Context-Aware Thought Visualization
by: Mehmood, Tariq, et al.
Published: (2025)
by: Mehmood, Tariq, et al.
Published: (2025)
DGRNet: Disagreement-Guided Refinement for Uncertainty-Aware Brain Tumor Segmentation
by: Mohammadi, Bahram, et al.
Published: (2026)
by: Mohammadi, Bahram, et al.
Published: (2026)
Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention
by: Khan, Zohaib, et al.
Published: (2024)
by: Khan, Zohaib, et al.
Published: (2024)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
by: Lin, Jingyang, et al.
Published: (2023)
by: Lin, Jingyang, et al.
Published: (2023)
Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
by: Xu, Zhongxing, et al.
Published: (2026)
by: Xu, Zhongxing, et al.
Published: (2026)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
Depth-Aware Image and Video Orientation Estimation
by: Alam, Muhammad Z., et al.
Published: (2026)
by: Alam, Muhammad Z., et al.
Published: (2026)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
by: Chen, Chao, et al.
Published: (2023)
by: Chen, Chao, et al.
Published: (2023)
Large Model based Sequential Keyframe Extraction for Video Summarization
by: Tan, Kailong, et al.
Published: (2024)
by: Tan, Kailong, et al.
Published: (2024)
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
Context-Aware Detection of Mixed Critical Events using Video Classification
by: Akhlaq, Filza, et al.
Published: (2024)
by: Akhlaq, Filza, et al.
Published: (2024)
Similar Items
-
CSTA: CNN-based Spatiotemporal Attention for Video Summarization
by: Son, Jaewon, et al.
Published: (2024) -
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024) -
Scaling Up Video Summarization Pretraining with Large Language Models
by: Argaw, Dawit Mureja, et al.
Published: (2024) -
EdgeDAM: Real-time Object Tracking for Mobile Devices
by: Raza, Syed Muhammad, et al.
Published: (2026) -
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
by: Ali, Muhammad, et al.
Published: (2024)