Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhiwang, Xu, Dong, Ouyang, Wanli, Tan, Chuanqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
VideoSAGE: Video Summarization with Graph Representation Learning
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
Comparing Learning Paradigms for Egocentric Video Summarization
by: Wen, Daniel
Published: (2025)
by: Wen, Daniel
Published: (2025)
An Integrated Framework for Multi-Granular Explanation of Video Summarization
by: Tsigos, Konstantinos, et al.
Published: (2024)
by: Tsigos, Konstantinos, et al.
Published: (2024)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Wolf: Dense Video Captioning with a World Summarization Framework
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
by: Wu, Yuanli, et al.
Published: (2025)
by: Wu, Yuanli, et al.
Published: (2025)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
by: Kim, Ye-Chan, et al.
Published: (2026)
by: Kim, Ye-Chan, et al.
Published: (2026)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
by: Eleftheriadis, Thomas, et al.
Published: (2025)
by: Eleftheriadis, Thomas, et al.
Published: (2025)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
by: Rajpal, Shreya, et al.
Published: (2026)
by: Rajpal, Shreya, et al.
Published: (2026)
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024)
by: Samel, Karan, et al.
Published: (2024)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
by: Barbakos, Spyros, et al.
Published: (2025)
by: Barbakos, Spyros, et al.
Published: (2025)
EdgeVidSum: Real-Time Personalized Video Summarization at the Edge
by: Mujtaba, Ghulam, et al.
Published: (2025)
by: Mujtaba, Ghulam, et al.
Published: (2025)
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
Model Compression using Progressive Channel Pruning
by: Guo, Jinyang, et al.
Published: (2025)
by: Guo, Jinyang, et al.
Published: (2025)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
by: Mei, Yuting, et al.
Published: (2024)
by: Mei, Yuting, et al.
Published: (2024)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
by: Díaz-Juan, Artur, et al.
Published: (2025)
by: Díaz-Juan, Artur, et al.
Published: (2025)
SD-VSum: A Method and Dataset for Script-Driven Video Summarization
by: Mylonas, Manolis, et al.
Published: (2025)
by: Mylonas, Manolis, et al.
Published: (2025)
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
by: Sarto, Sara, et al.
Published: (2024)
by: Sarto, Sara, et al.
Published: (2024)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
by: Li, Zhiyuan, et al.
Published: (2023)
by: Li, Zhiyuan, et al.
Published: (2023)
Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation
by: Avogaro, Niccolo, et al.
Published: (2025)
by: Avogaro, Niccolo, et al.
Published: (2025)
Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos
by: Das, Sarmistha, et al.
Published: (2025)
by: Das, Sarmistha, et al.
Published: (2025)
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
by: Jung, Woojun, et al.
Published: (2026)
by: Jung, Woojun, et al.
Published: (2026)
LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
by: Patel, Alkesh, et al.
Published: (2026)
by: Patel, Alkesh, et al.
Published: (2026)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026)
by: Kim, Sumin, et al.
Published: (2026)
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
by: Liu, Dongqi, et al.
Published: (2025)
by: Liu, Dongqi, et al.
Published: (2025)
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023)
by: Shen, Yaojie, et al.
Published: (2023)
VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
by: Jiang, Pengfei, et al.
Published: (2025)
by: Jiang, Pengfei, et al.
Published: (2025)
Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
HierSum: A Global and Local Attention Mechanism for Video Summarization
by: Beedu, Apoorva, et al.
Published: (2025)
by: Beedu, Apoorva, et al.
Published: (2025)
Similar Items
-
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025) -
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023) -
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
by: Zhang, Yan, et al.
Published: (2025) -
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024) -
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)