Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Zirui, Isobe, Takashi, Shen, Tong, Jia, Xu, Zhao, Jianbin, Li, Xiaomin, Ge, Mengmeng, Li, Baolu, Wang, Qinghe, Li, Dong, Zhou, Dong, Zhuge, Yunzhi, Lu, Huchuan, Barsoum, Emad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AMD-Hummingbird: Towards an Efficient Text-to-Video Model
di: Isobe, Takashi, et al.
Pubblicazione: (2025)
di: Isobe, Takashi, et al.
Pubblicazione: (2025)
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
di: Ge, Mengmeng, et al.
Pubblicazione: (2026)
di: Ge, Mengmeng, et al.
Pubblicazione: (2026)
StableIdentity: Inserting Anybody into Anywhere at First Sight
di: Wang, Qinghe, et al.
Pubblicazione: (2024)
di: Wang, Qinghe, et al.
Pubblicazione: (2024)
ReNeg: Learning Negative Embedding with Reward Guidance
di: Li, Xiaomin, et al.
Pubblicazione: (2024)
di: Li, Xiaomin, et al.
Pubblicazione: (2024)
CharacterFactory: Sampling Consistent Characters with GANs for Diffusion Models
di: Wang, Qinghe, et al.
Pubblicazione: (2024)
di: Wang, Qinghe, et al.
Pubblicazione: (2024)
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
di: Li, Xiaomin, et al.
Pubblicazione: (2024)
di: Li, Xiaomin, et al.
Pubblicazione: (2024)
VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
di: Li, Baolu, et al.
Pubblicazione: (2025)
di: Li, Baolu, et al.
Pubblicazione: (2025)
E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
di: Shen, Tong, et al.
Pubblicazione: (2025)
di: Shen, Tong, et al.
Pubblicazione: (2025)
LADDER: An Efficient Framework for Video Frame Interpolation
di: Shen, Tong, et al.
Pubblicazione: (2024)
di: Shen, Tong, et al.
Pubblicazione: (2024)
Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios
di: Li, Xiaomin, et al.
Pubblicazione: (2026)
di: Li, Xiaomin, et al.
Pubblicazione: (2026)
Edit as You See: Image-guided Video Editing via Masked Motion Modeling
di: Huang, Zhi-Lin, et al.
Pubblicazione: (2025)
di: Huang, Zhi-Lin, et al.
Pubblicazione: (2025)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
di: Gong, Sitong, et al.
Pubblicazione: (2025)
di: Gong, Sitong, et al.
Pubblicazione: (2025)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
MonoGS++: Fast and Accurate Monocular RGB Gaussian SLAM
di: Li, Renwu, et al.
Pubblicazione: (2025)
di: Li, Renwu, et al.
Pubblicazione: (2025)
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
di: Zhuge, Yunzhi, et al.
Pubblicazione: (2025)
di: Zhuge, Yunzhi, et al.
Pubblicazione: (2025)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
di: Xiong, Haomiao, et al.
Pubblicazione: (2025)
di: Xiong, Haomiao, et al.
Pubblicazione: (2025)
Complementary and Contrastive Learning for Audio-Visual Segmentation
di: Gong, Sitong, et al.
Pubblicazione: (2025)
di: Gong, Sitong, et al.
Pubblicazione: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
di: Li, Guanchen, et al.
Pubblicazione: (2025)
di: Li, Guanchen, et al.
Pubblicazione: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
di: Li, Zekai, et al.
Pubblicazione: (2026)
di: Li, Zekai, et al.
Pubblicazione: (2026)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
di: Xu, Yixing, et al.
Pubblicazione: (2025)
di: Xu, Yixing, et al.
Pubblicazione: (2025)
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
di: Wang, Qinghe, et al.
Pubblicazione: (2025)
di: Wang, Qinghe, et al.
Pubblicazione: (2025)
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
di: Liu, Qing'an, et al.
Pubblicazione: (2026)
di: Liu, Qing'an, et al.
Pubblicazione: (2026)
Parameter Aware Mamba Model for Multi-task Dense Prediction
di: Yu, Xinzhuo, et al.
Pubblicazione: (2025)
di: Yu, Xinzhuo, et al.
Pubblicazione: (2025)
Learning Universal Features for Generalizable Image Forgery Localization
di: Zhao, Hengrun, et al.
Pubblicazione: (2025)
di: Zhao, Hengrun, et al.
Pubblicazione: (2025)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
di: Gong, Sitong, et al.
Pubblicazione: (2025)
di: Gong, Sitong, et al.
Pubblicazione: (2025)
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
di: Zhang, Wenbo, et al.
Pubblicazione: (2024)
di: Zhang, Wenbo, et al.
Pubblicazione: (2024)
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
di: Feng, Xiyan, et al.
Pubblicazione: (2026)
di: Feng, Xiyan, et al.
Pubblicazione: (2026)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
di: Li, Zhe, et al.
Pubblicazione: (2025)
di: Li, Zhe, et al.
Pubblicazione: (2025)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
di: An, Zihao, et al.
Pubblicazione: (2026)
di: An, Zihao, et al.
Pubblicazione: (2026)
DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting
di: Yang, Yicheng, et al.
Pubblicazione: (2024)
di: Yang, Yicheng, et al.
Pubblicazione: (2024)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
di: Wang, Shuai, et al.
Pubblicazione: (2025)
di: Wang, Shuai, et al.
Pubblicazione: (2025)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
di: Xu, Yixing, et al.
Pubblicazione: (2025)
di: Xu, Yixing, et al.
Pubblicazione: (2025)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
di: Xiong, Haomiao, et al.
Pubblicazione: (2025)
di: Xiong, Haomiao, et al.
Pubblicazione: (2025)
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
di: Zhang, Lu, et al.
Pubblicazione: (2025)
di: Zhang, Lu, et al.
Pubblicazione: (2025)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
di: Gong, Sitong, et al.
Pubblicazione: (2025)
di: Gong, Sitong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AMD-Hummingbird: Towards an Efficient Text-to-Video Model
di: Isobe, Takashi, et al.
Pubblicazione: (2025) -
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
di: Ge, Mengmeng, et al.
Pubblicazione: (2026) -
StableIdentity: Inserting Anybody into Anywhere at First Sight
di: Wang, Qinghe, et al.
Pubblicazione: (2024) -
ReNeg: Learning Negative Embedding with Reward Guidance
di: Li, Xiaomin, et al.
Pubblicazione: (2024) -
CharacterFactory: Sampling Consistent Characters with GANs for Diffusion Models
di: Wang, Qinghe, et al.
Pubblicazione: (2024)