Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kothandaraman, Divya, Sohn, Kihyuk, Villegas, Ruben, Voigtlaender, Paul, Manocha, Dinesh, Babaeizadeh, Mohammad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Financial Models in Generative Art: Black-Scholes-Inspired Concept Blending in Text-to-Image Diffusion
by: Kothandaraman, Divya, et al.
Published: (2024)
by: Kothandaraman, Divya, et al.
Published: (2024)
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
by: Kothandaraman, Divya, et al.
Published: (2022)
by: Kothandaraman, Divya, et al.
Published: (2022)
SS-SFDA : Self-Supervised Source-Free Domain Adaptation for Road Segmentation in Hazardous Environments
by: Kothandaraman, Divya, et al.
Published: (2020)
by: Kothandaraman, Divya, et al.
Published: (2020)
BoMuDANet: Unsupervised Adaptation for Visual Scene Understanding in Unstructured Driving Environments
by: Kothandaraman, Divya, et al.
Published: (2020)
by: Kothandaraman, Divya, et al.
Published: (2020)
HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View
by: Kothandaraman, Divya, et al.
Published: (2023)
by: Kothandaraman, Divya, et al.
Published: (2023)
ImPoster: Text and Frequency Guidance for Subject Driven Action Personalization using Diffusion Models
by: Kothandaraman, Divya, et al.
Published: (2024)
by: Kothandaraman, Divya, et al.
Published: (2024)
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
3D-free meets 3D priors: Novel View Synthesis from a Single Image with Pretrained Diffusion Guidance
by: Kang, Taewon, et al.
Published: (2024)
by: Kang, Taewon, et al.
Published: (2024)
Placing Human Animations into 3D Scenes by Learning Interaction- and Geometry-Driven Keyframes
by: Mullen Jr, James F., et al.
Published: (2022)
by: Mullen Jr, James F., et al.
Published: (2022)
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
Zero-Shot Personalized Camera Motion Control for Image-to-Video Synthesis
by: Guhan, Pooja, et al.
Published: (2025)
by: Guhan, Pooja, et al.
Published: (2025)
SALAD: Source-free Active Label-Agnostic Domain Adaptation for Classification, Segmentation and Detection
by: Kothandaraman, Divya, et al.
Published: (2022)
by: Kothandaraman, Divya, et al.
Published: (2022)
DreamFlow: High-Quality Text-to-3D Generation by Approximating Probability Flow
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
by: Lee, Yonghan, et al.
Published: (2026)
by: Lee, Yonghan, et al.
Published: (2026)
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
Beyond Memorization: Selective Learning for Copyright-Safe Diffusion Model Training
by: Kothandaraman, Divya, et al.
Published: (2025)
by: Kothandaraman, Divya, et al.
Published: (2025)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Point-VOS: Pointing Up Video Object Segmentation
by: Zulfikar, Idil Esen, et al.
Published: (2024)
by: Zulfikar, Idil Esen, et al.
Published: (2024)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
by: Lokegaonkar, Vaibhavi, et al.
Published: (2026)
by: Lokegaonkar, Vaibhavi, et al.
Published: (2026)
Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention
by: Bhattacharya, Uttaran, et al.
Published: (2022)
by: Bhattacharya, Uttaran, et al.
Published: (2022)
SyncTrack4D: Cross-Video Motion Alignment and Video Synchronization for Multi-Video 4D Gaussian Splatting
by: Lee, Yonghan, et al.
Published: (2025)
by: Lee, Yonghan, et al.
Published: (2025)
FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
by: Ding, Ganggui, et al.
Published: (2024)
by: Ding, Ganggui, et al.
Published: (2024)
PACE: Data-Driven Virtual Agent Interaction in Dense and Cluttered Environments
by: Mullen, James, et al.
Published: (2023)
by: Mullen, James, et al.
Published: (2023)
Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion
by: Tang, Zhenggang, et al.
Published: (2026)
by: Tang, Zhenggang, et al.
Published: (2026)
RegionRoute: Regional Style Transfer with Diffusion Model
by: Chen, Bowen, et al.
Published: (2026)
by: Chen, Bowen, et al.
Published: (2026)
Low-Bitrate Video Compression through Semantic-Conditioned Diffusion
by: Wang, Lingdong, et al.
Published: (2025)
by: Wang, Lingdong, et al.
Published: (2025)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
by: Chen, Hong, et al.
Published: (2024)
by: Chen, Hong, et al.
Published: (2024)
HighlightMe: Detecting Highlights from Human-Centric Videos
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
User-Friendly Customized Generation with Multi-Modal Prompts
by: Zhong, Linhao, et al.
Published: (2024)
by: Zhong, Linhao, et al.
Published: (2024)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
by: Huang, Yuzhou, et al.
Published: (2025)
by: Huang, Yuzhou, et al.
Published: (2025)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models
by: Chen, Hong, et al.
Published: (2023)
by: Chen, Hong, et al.
Published: (2023)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
by: Bhattacharya, Uttaran, et al.
Published: (2024)
by: Bhattacharya, Uttaran, et al.
Published: (2024)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
by: Ren, Yixuan, et al.
Published: (2024)
by: Ren, Yixuan, et al.
Published: (2024)
BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation
by: Hu, Panwen, et al.
Published: (2025)
by: Hu, Panwen, et al.
Published: (2025)
DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
by: Wu, Fangtai, et al.
Published: (2025)
by: Wu, Fangtai, et al.
Published: (2025)
Similar Items
-
Financial Models in Generative Art: Black-Scholes-Inspired Concept Blending in Text-to-Image Diffusion
by: Kothandaraman, Divya, et al.
Published: (2024) -
Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition
by: Kothandaraman, Divya, et al.
Published: (2022) -
SS-SFDA : Self-Supervised Source-Free Domain Adaptation for Road Segmentation in Hazardous Environments
by: Kothandaraman, Divya, et al.
Published: (2020) -
BoMuDANet: Unsupervised Adaptation for Visual Scene Understanding in Unstructured Driving Environments
by: Kothandaraman, Divya, et al.
Published: (2020) -
HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View
by: Kothandaraman, Divya, et al.
Published: (2023)