DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Weijie, Liu, Mushui, Yu, Yunlong, Wang, Zhao, Wu, Chao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
di: Wu, Fangtai, et al.
Pubblicazione: (2025)
di: Wu, Fangtai, et al.
Pubblicazione: (2025)
Hybrid Mask Generation for Infrared Small Target Detection with Single-Point Supervision
di: He, Weijie, et al.
Pubblicazione: (2024)
di: He, Weijie, et al.
Pubblicazione: (2024)
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
di: Seitzer, Maximilian, et al.
Pubblicazione: (2023)
di: Seitzer, Maximilian, et al.
Pubblicazione: (2023)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
Fully Fine-tuned CLIP Models are Efficient Few-Shot Learners
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
di: Li, Bozheng, et al.
Pubblicazione: (2024)
di: Li, Bozheng, et al.
Pubblicazione: (2024)
LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
Multitwine: Multi-Object Compositing with Text and Layout Control
di: Tarrés, Gemma Canet, et al.
Pubblicazione: (2025)
di: Tarrés, Gemma Canet, et al.
Pubblicazione: (2025)
LAYOUTDREAMER: Physics-guided Layout for Text-to-3D Compositional Scene Generation
di: Zhou, Yang, et al.
Pubblicazione: (2025)
di: Zhou, Yang, et al.
Pubblicazione: (2025)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
Envisioning Class Entity Reasoning by Large Language Models for Few-shot Learning
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
di: Wu, Yuxuan, et al.
Pubblicazione: (2025)
di: Wu, Yuxuan, et al.
Pubblicazione: (2025)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
di: She, D., et al.
Pubblicazione: (2025)
di: She, D., et al.
Pubblicazione: (2025)
CM-UNet: Hybrid CNN-Mamba UNet for Remote Sensing Image Semantic Segmentation
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
di: Horita, Daichi, et al.
Pubblicazione: (2023)
di: Horita, Daichi, et al.
Pubblicazione: (2023)
PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language Models
di: He, Runze, et al.
Pubblicazione: (2025)
di: He, Runze, et al.
Pubblicazione: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
di: Liu, Xiangrui, et al.
Pubblicazione: (2025)
di: Liu, Xiangrui, et al.
Pubblicazione: (2025)
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
di: Ye, Bo, et al.
Pubblicazione: (2026)
di: Ye, Bo, et al.
Pubblicazione: (2026)
RestorerID: Towards Tuning-Free Face Restoration with ID Preservation
di: Ying, Jiacheng, et al.
Pubblicazione: (2024)
di: Ying, Jiacheng, et al.
Pubblicazione: (2024)
VideoTetris: Towards Compositional Text-to-Video Generation
di: Tian, Ye, et al.
Pubblicazione: (2024)
di: Tian, Ye, et al.
Pubblicazione: (2024)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
di: Feng, Weixi, et al.
Pubblicazione: (2025)
di: Feng, Weixi, et al.
Pubblicazione: (2025)
TableSeq: Unified Generation of Structure, Content, and Layout
di: Hamdi, Laziz, et al.
Pubblicazione: (2026)
di: Hamdi, Laziz, et al.
Pubblicazione: (2026)
UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation
di: Liu, Zeyang, et al.
Pubblicazione: (2025)
di: Liu, Zeyang, et al.
Pubblicazione: (2025)
Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation
di: Zhou, Junwei, et al.
Pubblicazione: (2026)
di: Zhou, Junwei, et al.
Pubblicazione: (2026)
Text-Animator: Controllable Visual Text Video Generation
di: Liu, Lin, et al.
Pubblicazione: (2024)
di: Liu, Lin, et al.
Pubblicazione: (2024)
DyCoRM: Dynamic Criterion-Aware Reward Modeling for Text-to-Image Generation
di: Qian, Jiaying, et al.
Pubblicazione: (2026)
di: Qian, Jiaying, et al.
Pubblicazione: (2026)
CSGO: Content-Style Composition in Text-to-Image Generation
di: Xing, Peng, et al.
Pubblicazione: (2024)
di: Xing, Peng, et al.
Pubblicazione: (2024)
SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior
di: Wang, Haoran, et al.
Pubblicazione: (2025)
di: Wang, Haoran, et al.
Pubblicazione: (2025)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
di: Huang, Kaiyi, et al.
Pubblicazione: (2024)
di: Huang, Kaiyi, et al.
Pubblicazione: (2024)
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
di: Zheng, Guangcong, et al.
Pubblicazione: (2023)
di: Zheng, Guangcong, et al.
Pubblicazione: (2023)
Generating Animated Layouts as Structured Text Representations
di: Shin, Yeonsang, et al.
Pubblicazione: (2025)
di: Shin, Yeonsang, et al.
Pubblicazione: (2025)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
di: Liu, Jinlong, et al.
Pubblicazione: (2026)
di: Liu, Jinlong, et al.
Pubblicazione: (2026)
MambaVSR: Content-Aware Scanning State Space Model for Video Super-Resolution
di: He, Linfeng, et al.
Pubblicazione: (2025)
di: He, Linfeng, et al.
Pubblicazione: (2025)
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
di: Shu, Yan, et al.
Pubblicazione: (2024)
di: Shu, Yan, et al.
Pubblicazione: (2024)
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
di: Cheng, Bo, et al.
Pubblicazione: (2024)
di: Cheng, Bo, et al.
Pubblicazione: (2024)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
di: Zhang, Hui, et al.
Pubblicazione: (2024)
di: Zhang, Hui, et al.
Pubblicazione: (2024)
Generating Synthetic Invoices via Layout-Preserving Content Replacement
di: V, Bevin, et al.
Pubblicazione: (2025)
di: V, Bevin, et al.
Pubblicazione: (2025)
LTOS: Layout-controllable Text-Object Synthesis via Adaptive Cross-attention Fusions
di: Zhao, Xiaoran, et al.
Pubblicazione: (2024)
di: Zhao, Xiaoran, et al.
Pubblicazione: (2024)
Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints
di: Fang, Chuan, et al.
Pubblicazione: (2023)
di: Fang, Chuan, et al.
Pubblicazione: (2023)
CoMo: Compositional Motion Customization for Text-to-Video Generation
di: Xu, Youcan, et al.
Pubblicazione: (2025)
di: Xu, Youcan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
di: Wu, Fangtai, et al.
Pubblicazione: (2025) -
Hybrid Mask Generation for Infrared Small Target Detection with Single-Point Supervision
di: He, Weijie, et al.
Pubblicazione: (2024) -
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
di: Seitzer, Maximilian, et al.
Pubblicazione: (2023) -
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
di: Liu, Mushui, et al.
Pubblicazione: (2024) -
Fully Fine-tuned CLIP Models are Efficient Few-Shot Learners
di: Liu, Mushui, et al.
Pubblicazione: (2024)