GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Eileen, Han, Caren, Poon, Josiah |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
von: Lin, Shih-Yao, et al.
Veröffentlicht: (2025)
von: Lin, Shih-Yao, et al.
Veröffentlicht: (2025)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
MSG-Chart: Multimodal Scene Graph for ChartQA
von: Dai, Yue, et al.
Veröffentlicht: (2024)
von: Dai, Yue, et al.
Veröffentlicht: (2024)
UVG-VPC: Voxelized Point Cloud Dataset for Visual Volumetric Video-based Coding
von: Gautier, Guillaume, et al.
Veröffentlicht: (2025)
von: Gautier, Guillaume, et al.
Veröffentlicht: (2025)
Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
von: Wang, Mengzhao, et al.
Veröffentlicht: (2024)
von: Wang, Mengzhao, et al.
Veröffentlicht: (2024)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2025)
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2025)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding
von: Zhong, Yaokun, et al.
Veröffentlicht: (2025)
von: Zhong, Yaokun, et al.
Veröffentlicht: (2025)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
von: Zhou, Kaichen, et al.
Veröffentlicht: (2026)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2026)
MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning
von: Xiong, Huiyu, et al.
Veröffentlicht: (2024)
von: Xiong, Huiyu, et al.
Veröffentlicht: (2024)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
von: Meng, Desen, et al.
Veröffentlicht: (2025)
von: Meng, Desen, et al.
Veröffentlicht: (2025)
Paragraph-to-Image Generation with Information-Enriched Diffusion Model
von: Wu, Weijia, et al.
Veröffentlicht: (2023)
von: Wu, Weijia, et al.
Veröffentlicht: (2023)
Live Video Captioning
von: Blanco-Fernández, Eduardo, et al.
Veröffentlicht: (2024)
von: Blanco-Fernández, Eduardo, et al.
Veröffentlicht: (2024)
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
Addressing the ID-Matching Challenge in Long Video Captioning
von: Yang, Zhantao, et al.
Veröffentlicht: (2025)
von: Yang, Zhantao, et al.
Veröffentlicht: (2025)
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
von: Shoman, Maged, et al.
Veröffentlicht: (2024)
von: Shoman, Maged, et al.
Veröffentlicht: (2024)
Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
Streaming Dense Video Captioning
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
von: Zhou, Xingyi, et al.
Veröffentlicht: (2024)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
Zero-Shot Paragraph-level Handwriting Imitation with Latent Diffusion Models
von: Mayr, Martin, et al.
Veröffentlicht: (2024)
von: Mayr, Martin, et al.
Veröffentlicht: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
MV-CC: Mask Enhanced Video Model for Remote Sensing Change Caption
von: Liu, Ruixun, et al.
Veröffentlicht: (2024)
von: Liu, Ruixun, et al.
Veröffentlicht: (2024)
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
von: Yu, Linhao, et al.
Veröffentlicht: (2025)
von: Yu, Linhao, et al.
Veröffentlicht: (2025)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
ViDiC: Video Difference Captioning
von: Wu, Jiangtao, et al.
Veröffentlicht: (2025)
von: Wu, Jiangtao, et al.
Veröffentlicht: (2025)
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
von: Xi, Zeyu, et al.
Veröffentlicht: (2025)
von: Xi, Zeyu, et al.
Veröffentlicht: (2025)
Progress-Aware Video Frame Captioning
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
von: Wang, Eileen, et al.
Veröffentlicht: (2024) -
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
von: Chen, Sishuo, et al.
Veröffentlicht: (2024) -
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2025) -
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
von: Lin, Shih-Yao, et al.
Veröffentlicht: (2025) -
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
von: Ding, Yihao, et al.
Veröffentlicht: (2024)