Multimodal Chaptering for Long-Form TV Newscast Video
Fuente:
arXiv
Saved in:
| Main Authors: | Guetari, Khalil, Tevissen, Yannis, Petitpont, Frédéric |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Retrieval Augmented Generation over Large Video Libraries
by: Tevissen, Yannis, et al.
Published: (2024)
by: Tevissen, Yannis, et al.
Published: (2024)
Inserting Faces inside Captions: Image Captioning with Attention Guided Merging
by: Tevissen, Yannis, et al.
Published: (2024)
by: Tevissen, Yannis, et al.
Published: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
by: Li, Jie, et al.
Published: (2025)
by: Li, Jie, et al.
Published: (2025)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
by: Liu, Jiajun, et al.
Published: (2024)
by: Liu, Jiajun, et al.
Published: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
by: Yu, Jiashuo, et al.
Published: (2025)
by: Yu, Jiashuo, et al.
Published: (2025)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
by: Zeng, Xiangyu, et al.
Published: (2024)
by: Zeng, Xiangyu, et al.
Published: (2024)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
by: Wang, Yihao, et al.
Published: (2024)
by: Wang, Yihao, et al.
Published: (2024)
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
by: Zhang, Yinghui, et al.
Published: (2025)
by: Zhang, Yinghui, et al.
Published: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
by: Breuer, Adam, et al.
Published: (2025)
by: Breuer, Adam, et al.
Published: (2025)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
by: Hao, Haoran, et al.
Published: (2025)
by: Hao, Haoran, et al.
Published: (2025)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
Video Seal: Open and Efficient Video Watermarking
by: Fernandez, Pierre, et al.
Published: (2024)
by: Fernandez, Pierre, et al.
Published: (2024)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
by: Wang, Zhenyu, et al.
Published: (2026)
by: Wang, Zhenyu, et al.
Published: (2026)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
by: Tao, Ziyuan, et al.
Published: (2025)
by: Tao, Ziyuan, et al.
Published: (2025)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
by: Gu, Jing, et al.
Published: (2024)
by: Gu, Jing, et al.
Published: (2024)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
by: Qing, Yuan, et al.
Published: (2026)
by: Qing, Yuan, et al.
Published: (2026)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations
by: Galarnyk, Michael, et al.
Published: (2025)
by: Galarnyk, Michael, et al.
Published: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Diagnosing and Re-learning for Balanced Multimodal Learning
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
by: Ku, Max, et al.
Published: (2024)
by: Ku, Max, et al.
Published: (2024)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Very Efficient Listwise Multimodal Reranking for Long Documents
by: Sun, Yiqun, et al.
Published: (2026)
by: Sun, Yiqun, et al.
Published: (2026)
Question-Answering Dense Video Events
by: Qin, Hangyu, et al.
Published: (2024)
by: Qin, Hangyu, et al.
Published: (2024)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
by: Ku, Max, et al.
Published: (2025)
by: Ku, Max, et al.
Published: (2025)
Multimodal Markup Document Models for Graphic Design Completion
by: Kikuchi, Kotaro, et al.
Published: (2024)
by: Kikuchi, Kotaro, et al.
Published: (2024)
Bernini: Latent Semantic Planning for Video Diffusion
by: Bernini Team, et al.
Published: (2026)
by: Bernini Team, et al.
Published: (2026)
Interactive Video Generation via Domain Adaptation
by: Rawal, Ishaan, et al.
Published: (2025)
by: Rawal, Ishaan, et al.
Published: (2025)
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
by: Phung, Thu Hang, et al.
Published: (2026)
by: Phung, Thu Hang, et al.
Published: (2026)
Image Conductor: Precision Control for Interactive Video Synthesis
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
Advance Fake Video Detection via Vision Transformers
by: Battocchio, Joy, et al.
Published: (2025)
by: Battocchio, Joy, et al.
Published: (2025)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
by: Barbakos, Spyros, et al.
Published: (2025)
by: Barbakos, Spyros, et al.
Published: (2025)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
by: Zhao, Ruixiang, et al.
Published: (2024)
by: Zhao, Ruixiang, et al.
Published: (2024)
Similar Items
-
Towards Retrieval Augmented Generation over Large Video Libraries
by: Tevissen, Yannis, et al.
Published: (2024) -
Inserting Faces inside Captions: Image Captioning with Attention Guided Merging
by: Tevissen, Yannis, et al.
Published: (2024) -
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025) -
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
by: Li, Jie, et al.
Published: (2025) -
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
by: Liu, Jiajun, et al.
Published: (2024)