MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhekai, Wang, Yuqing, Zhang, Manyuan, Liu, Xihui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
von: Chen, Zhekai, et al.
Veröffentlicht: (2025)
von: Chen, Zhekai, et al.
Veröffentlicht: (2025)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
von: Yu, Jiwen, et al.
Veröffentlicht: (2025)
von: Yu, Jiwen, et al.
Veröffentlicht: (2025)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
von: Yue, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Yue, Xiaoyu, et al.
Veröffentlicht: (2025)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
von: Teng, Yao, et al.
Veröffentlicht: (2025)
von: Teng, Yao, et al.
Veröffentlicht: (2025)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
von: Qiu, Lu, et al.
Veröffentlicht: (2025)
von: Qiu, Lu, et al.
Veröffentlicht: (2025)
4Diffusion: Multi-view Video Diffusion Model for 4D Generation
von: Zhang, Haiyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haiyu, et al.
Veröffentlicht: (2024)
GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
von: Chen, Zhekai, et al.
Veröffentlicht: (2024)
von: Chen, Zhekai, et al.
Veröffentlicht: (2024)
Personalized Text-to-Image Generation with Auto-Regressive Models
von: Sun, Kaiyue, et al.
Veröffentlicht: (2025)
von: Sun, Kaiyue, et al.
Veröffentlicht: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
von: Chu, Ruihang, et al.
Veröffentlicht: (2025)
von: Chu, Ruihang, et al.
Veröffentlicht: (2025)
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
von: Huang, Binyuan, et al.
Veröffentlicht: (2026)
von: Huang, Binyuan, et al.
Veröffentlicht: (2026)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
von: Ge, Junqi, et al.
Veröffentlicht: (2024)
MultiRef: Controllable Image Generation with Multiple Visual References
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
Gen-Searcher: Reinforcing Agentic Search for Image Generation
von: Feng, Kaituo, et al.
Veröffentlicht: (2026)
von: Feng, Kaituo, et al.
Veröffentlicht: (2026)
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
von: Sun, Kaiyue, et al.
Veröffentlicht: (2025)
von: Sun, Kaiyue, et al.
Veröffentlicht: (2025)
Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
von: Wang, Yunnan, et al.
Veröffentlicht: (2024)
von: Wang, Yunnan, et al.
Veröffentlicht: (2024)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
von: Huang, Kaiyi, et al.
Veröffentlicht: (2024)
von: Huang, Kaiyi, et al.
Veröffentlicht: (2024)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
MultiWorld: Scalable Multi-Agent Multi-View Video World Models
von: Wu, Haoyu, et al.
Veröffentlicht: (2026)
von: Wu, Haoyu, et al.
Veröffentlicht: (2026)
Parallelized Autoregressive Visual Generation
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting
von: Liu, Taorong, et al.
Veröffentlicht: (2023)
von: Liu, Taorong, et al.
Veröffentlicht: (2023)
MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
Conditional Text-to-Image Generation with Reference Guidance
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
von: Chen, Haoyu, et al.
Veröffentlicht: (2026)
von: Chen, Haoyu, et al.
Veröffentlicht: (2026)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
von: Shang, Yu, et al.
Veröffentlicht: (2025)
von: Shang, Yu, et al.
Veröffentlicht: (2025)
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
von: Yan, Cilin, et al.
Veröffentlicht: (2025)
von: Yan, Cilin, et al.
Veröffentlicht: (2025)
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
von: Xiong, Tianwei, et al.
Veröffentlicht: (2025)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2025)
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
von: Guo, Qin, et al.
Veröffentlicht: (2025)
von: Guo, Qin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
von: Chen, Zhekai, et al.
Veröffentlicht: (2025) -
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
von: Yu, Jiwen, et al.
Veröffentlicht: (2025) -
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
von: Yue, Xiaoyu, et al.
Veröffentlicht: (2025) -
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2024) -
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
von: Teng, Yao, et al.
Veröffentlicht: (2025)