Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Abbasi, Mehryar, Hadizadeh, Hadi, Saeedi, Parvaneh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
di: Breuer, Adam, et al.
Pubblicazione: (2025)
di: Breuer, Adam, et al.
Pubblicazione: (2025)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
di: Barbakos, Spyros, et al.
Pubblicazione: (2025)
di: Barbakos, Spyros, et al.
Pubblicazione: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
STIV: Scalable Text and Image Conditioned Video Generation
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising
di: Singer, Assaf, et al.
Pubblicazione: (2025)
di: Singer, Assaf, et al.
Pubblicazione: (2025)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
di: Li, Yumeng, et al.
Pubblicazione: (2024)
di: Li, Yumeng, et al.
Pubblicazione: (2024)
Latent Space Probing for Adult Content Detection in Video Generative Models
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2025)
di: Li, Quanhao, et al.
Pubblicazione: (2025)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2026)
di: Li, Quanhao, et al.
Pubblicazione: (2026)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
di: Yu, Lijun
Pubblicazione: (2024)
di: Yu, Lijun
Pubblicazione: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
LayerT2V: A Unified Multi-Layer Video Generation Framework
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
Video-based Music Generation
di: Sulun, Serkan
Pubblicazione: (2026)
di: Sulun, Serkan
Pubblicazione: (2026)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
di: Girdhar, Rohit, et al.
Pubblicazione: (2023)
di: Girdhar, Rohit, et al.
Pubblicazione: (2023)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
Conditioning GAN Without Training Dataset
di: Mekonnen, Kidist Amde
Pubblicazione: (2024)
di: Mekonnen, Kidist Amde
Pubblicazione: (2024)
Diffusion Model-Based Video Editing: A Survey
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
di: Yang, Dejie, et al.
Pubblicazione: (2024)
di: Yang, Dejie, et al.
Pubblicazione: (2024)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
di: Lin, Ronghao, et al.
Pubblicazione: (2024)
di: Lin, Ronghao, et al.
Pubblicazione: (2024)
Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2022)
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2022)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
di: Muqeet, Abdul, et al.
Pubblicazione: (2023)
di: Muqeet, Abdul, et al.
Pubblicazione: (2023)
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2025)
di: Croitoru, Florinel-Alin, et al.
Pubblicazione: (2025)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
di: Wang, Qiang, et al.
Pubblicazione: (2025)
di: Wang, Qiang, et al.
Pubblicazione: (2025)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
Flow Generator Matching
di: Huang, Zemin, et al.
Pubblicazione: (2024)
di: Huang, Zemin, et al.
Pubblicazione: (2024)
Generating Illustrated Instructions
di: Menon, Sachit, et al.
Pubblicazione: (2023)
di: Menon, Sachit, et al.
Pubblicazione: (2023)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
di: Mei, Yuting, et al.
Pubblicazione: (2024)
di: Mei, Yuting, et al.
Pubblicazione: (2024)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
di: Díaz-Juan, Artur, et al.
Pubblicazione: (2025)
di: Díaz-Juan, Artur, et al.
Pubblicazione: (2025)
SD-VSum: A Method and Dataset for Script-Driven Video Summarization
di: Mylonas, Manolis, et al.
Pubblicazione: (2025)
di: Mylonas, Manolis, et al.
Pubblicazione: (2025)
Improving Visual Representation Alignment Generation with GRPO
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
On-the-fly Modulation for Balanced Multimodal Learning
di: Wei, Yake, et al.
Pubblicazione: (2024)
di: Wei, Yake, et al.
Pubblicazione: (2024)
A Systematic Review on Long-Tailed Learning
di: Zhang, Chongsheng, et al.
Pubblicazione: (2024)
di: Zhang, Chongsheng, et al.
Pubblicazione: (2024)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
di: Dong, Hao, et al.
Pubblicazione: (2026)
di: Dong, Hao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
di: Chen, Weifeng, et al.
Pubblicazione: (2023) -
Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012
di: Breuer, Adam, et al.
Pubblicazione: (2025) -
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
di: Barbakos, Spyros, et al.
Pubblicazione: (2025) -
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024) -
STIV: Scalable Text and Image Conditioned Video Generation
di: Lin, Zongyu, et al.
Pubblicazione: (2024)