TeaserGen: Generating Teasers for Long Documentaries
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Weihan, Liang, Paul Pu, Kim, Haven, McAuley, Julian, Berg-Kirkpatrick, Taylor, Dong, Hao-Wen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
by: Xu, Weihan, et al.
Published: (2024)
by: Xu, Weihan, et al.
Published: (2024)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
by: Kim, Haven, et al.
Published: (2025)
by: Kim, Haven, et al.
Published: (2025)
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
by: Kim, Haven, et al.
Published: (2026)
by: Kim, Haven, et al.
Published: (2026)
LVCHAT: Facilitating Long Video Comprehension
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
by: Kim, Haven, et al.
Published: (2026)
by: Kim, Haven, et al.
Published: (2026)
Imagery as Inquiry: Exploring A Multimodal Dataset for Conversational Recommendation
by: Yoon, Se-eun, et al.
Published: (2024)
by: Yoon, Se-eun, et al.
Published: (2024)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Classification of horospherical invariant measures in higher rank: Teaser
by: Choi, Inhyeok, et al.
Published: (2025)
by: Choi, Inhyeok, et al.
Published: (2025)
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
by: Yang, Qihui, et al.
Published: (2025)
by: Yang, Qihui, et al.
Published: (2025)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
by: Long, Phillip, et al.
Published: (2024)
by: Long, Phillip, et al.
Published: (2024)
LaViC: Adapting Large Vision-Language Models to Visually-Aware Conversational Recommendation
by: Jeon, Hyunsik, et al.
Published: (2025)
by: Jeon, Hyunsik, et al.
Published: (2025)
Steering Autoregressive Music Generation with Recursive Feature Machines
by: Zhao, Daniel, et al.
Published: (2025)
by: Zhao, Daniel, et al.
Published: (2025)
PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior
by: Wu, Junda, et al.
Published: (2025)
by: Wu, Junda, et al.
Published: (2025)
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
by: Gao, Xin, et al.
Published: (2026)
by: Gao, Xin, et al.
Published: (2026)
Optical Context Compression Is Just (Bad) Autoencoding
by: Lee, Ivan Yee, et al.
Published: (2025)
by: Lee, Ivan Yee, et al.
Published: (2025)
Are you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
by: Li, Xintong, et al.
Published: (2025)
by: Li, Xintong, et al.
Published: (2025)
A Fashion Item Recommendation Model in Hyperbolic Space
by: Shimizu, Ryotaro, et al.
Published: (2024)
by: Shimizu, Ryotaro, et al.
Published: (2024)
PodReels: Human-AI Co-Creation of Video Podcast Teasers
by: Wang, Sitong, et al.
Published: (2023)
by: Wang, Sitong, et al.
Published: (2023)
EmoWear: Exploring Emotional Teasers for Voice Message Interaction on Smartwatches
by: An, Pengcheng, et al.
Published: (2024)
by: An, Pengcheng, et al.
Published: (2024)
Alt-Text with Context: Improving Accessibility for Images on Twitter
by: Srivatsan, Nikita, et al.
Published: (2023)
by: Srivatsan, Nikita, et al.
Published: (2023)
Interactive Mars Image Content-Based Search with Interpretable Machine Learning
by: Vasu, Bhavan, et al.
Published: (2024)
by: Vasu, Bhavan, et al.
Published: (2024)
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
by: Huang, Jingyue, et al.
Published: (2025)
by: Huang, Jingyue, et al.
Published: (2025)
Expressiveness Limits of Autoregressive Semantic ID Generation in Generative Recommendation
by: Hou, Yupeng, et al.
Published: (2026)
by: Hou, Yupeng, et al.
Published: (2026)
Teaser - Upemba Key Landscape for Conservation Land Cover and Validation Datasets (2016)
by: Szantoi, Zoltan, et al.
Published: (2020)
by: Szantoi, Zoltan, et al.
Published: (2020)
A Near-Infrared and Millimeter Study of the Rosette Molecular Cloud. An EMIR Teaser
by: Carlos G. Román-Zúñiga
Published: (2005)
by: Carlos G. Román-Zúñiga
Published: (2005)
Teaser - Takamanda Key Landscape for Conservation Land Cover and Validation Datasets (2016)
by: Szantoi, Zoltan, et al.
Published: (2020)
by: Szantoi, Zoltan, et al.
Published: (2020)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
by: Wang, Chuhan, et al.
Published: (2026)
by: Wang, Chuhan, et al.
Published: (2026)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
by: Chen, Danlu, et al.
Published: (2024)
by: Chen, Danlu, et al.
Published: (2024)
TokensGen: Harnessing Condensed Tokens for Long Video Generation
by: Ouyang, Wenqi, et al.
Published: (2025)
by: Ouyang, Wenqi, et al.
Published: (2025)
Teaser - The Great Limpopo Key Landscape for Conservation Land Cover and Validation Datasets (2016)
by: Szantoi, Zoltan, et al.
Published: (2020)
by: Szantoi, Zoltan, et al.
Published: (2020)
Teaser - Tai-Sapo Key Landscape for Conservation Land Cover and Validation Datasets (2016)
by: Szantoi, Zoltan, et al.
Published: (2020)
by: Szantoi, Zoltan, et al.
Published: (2020)
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
by: Wang, Shijian, et al.
Published: (2024)
by: Wang, Shijian, et al.
Published: (2024)
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
by: Wu, Junda, et al.
Published: (2025)
by: Wu, Junda, et al.
Published: (2025)
Progressive Compositionality in Text-to-Image Generative Models
by: Han, Evans Xu, et al.
Published: (2024)
by: Han, Evans Xu, et al.
Published: (2024)
Similar Items
-
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
by: Xu, Weihan, et al.
Published: (2025) -
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
by: Xu, Weihan, et al.
Published: (2024) -
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
by: Kim, Haven, et al.
Published: (2025) -
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
by: Kim, Haven, et al.
Published: (2026) -
LVCHAT: Facilitating Long Video Comprehension
by: Wang, Yu, et al.
Published: (2024)