Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yabo, Li, Kunchang, Zhou, Dewei, Huang, Xinyu, Wang, Xun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
por: Liao, Chao, et al.
Publicado: (2025)
por: Liao, Chao, et al.
Publicado: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
por: Li, Kunchang, et al.
Publicado: (2022)
por: Li, Kunchang, et al.
Publicado: (2022)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
por: Zhou, Chao, et al.
Publicado: (2025)
por: Zhou, Chao, et al.
Publicado: (2025)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
por: Liu, Jie, et al.
Publicado: (2026)
por: Liu, Jie, et al.
Publicado: (2026)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
por: Chen, Junyi, et al.
Publicado: (2026)
por: Chen, Junyi, et al.
Publicado: (2026)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
por: Bu, Wendong, et al.
Publicado: (2025)
por: Bu, Wendong, et al.
Publicado: (2025)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
por: Chen, Dongping, et al.
Publicado: (2024)
por: Chen, Dongping, et al.
Publicado: (2024)
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
por: Liu, Qingyang, et al.
Publicado: (2026)
por: Liu, Qingyang, et al.
Publicado: (2026)
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
por: Nie, Ming, et al.
Publicado: (2026)
por: Nie, Ming, et al.
Publicado: (2026)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
por: Zhou, Dewei, et al.
Publicado: (2025)
por: Zhou, Dewei, et al.
Publicado: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
por: Tian, Changyao, et al.
Publicado: (2024)
por: Tian, Changyao, et al.
Publicado: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
por: Zhou, Dewei, et al.
Publicado: (2024)
por: Zhou, Dewei, et al.
Publicado: (2024)
Multi-Sentence Grounding for Long-term Instructional Video
por: Li, Zeqian, et al.
Publicado: (2023)
por: Li, Zeqian, et al.
Publicado: (2023)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
por: Lin, Weifeng, et al.
Publicado: (2024)
por: Lin, Weifeng, et al.
Publicado: (2024)
MANTIS: Interleaved Multi-Image Instruction Tuning
por: Jiang, Dongfu, et al.
Publicado: (2024)
por: Jiang, Dongfu, et al.
Publicado: (2024)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
por: Li, Yanlin, et al.
Publicado: (2026)
por: Li, Yanlin, et al.
Publicado: (2026)
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
por: Pu, Yujiang, et al.
Publicado: (2025)
por: Pu, Yujiang, et al.
Publicado: (2025)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
por: Chen, Haoyu, et al.
Publicado: (2026)
por: Chen, Haoyu, et al.
Publicado: (2026)
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
por: Zhou, Dewei, et al.
Publicado: (2024)
por: Zhou, Dewei, et al.
Publicado: (2024)
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
por: Yang, Jinrui, et al.
Publicado: (2026)
por: Yang, Jinrui, et al.
Publicado: (2026)
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
por: Zhou, Dewei, et al.
Publicado: (2024)
por: Zhou, Dewei, et al.
Publicado: (2024)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
por: Wang, Junke, et al.
Publicado: (2025)
por: Wang, Junke, et al.
Publicado: (2025)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
por: Zhang, Lei, et al.
Publicado: (2026)
por: Zhang, Lei, et al.
Publicado: (2026)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
por: Zhou, Dewei, et al.
Publicado: (2025)
por: Zhou, Dewei, et al.
Publicado: (2025)
Holistic Evaluation for Interleaved Text-and-Image Generation
por: Liu, Minqian, et al.
Publicado: (2024)
por: Liu, Minqian, et al.
Publicado: (2024)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
por: Li, Xinhao, et al.
Publicado: (2024)
por: Li, Xinhao, et al.
Publicado: (2024)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
por: Chen, Xinyan, et al.
Publicado: (2025)
por: Chen, Xinyan, et al.
Publicado: (2025)
ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
por: Wang, Lihong, et al.
Publicado: (2025)
por: Wang, Lihong, et al.
Publicado: (2025)
Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval
por: Du, Yongchao, et al.
Publicado: (2024)
por: Du, Yongchao, et al.
Publicado: (2024)
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
por: Chow, Wei, et al.
Publicado: (2025)
por: Chow, Wei, et al.
Publicado: (2025)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
por: Feng, Yukang, et al.
Publicado: (2025)
por: Feng, Yukang, et al.
Publicado: (2025)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
por: Zhou, Pengfei, et al.
Publicado: (2024)
por: Zhou, Pengfei, et al.
Publicado: (2024)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
por: Zhang, Huanyu, et al.
Publicado: (2026)
por: Zhang, Huanyu, et al.
Publicado: (2026)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
por: Li, Qingyun, et al.
Publicado: (2024)
por: Li, Qingyun, et al.
Publicado: (2024)
Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation
por: Lin, Zeteng, et al.
Publicado: (2025)
por: Lin, Zeteng, et al.
Publicado: (2025)
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
por: Hong, Lingyi, et al.
Publicado: (2026)
por: Hong, Lingyi, et al.
Publicado: (2026)
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
por: Yu, Haodong, et al.
Publicado: (2026)
por: Yu, Haodong, et al.
Publicado: (2026)
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
por: Dong, Shuai, et al.
Publicado: (2025)
por: Dong, Shuai, et al.
Publicado: (2025)
SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens
por: Zhang, Xiaoyan, et al.
Publicado: (2026)
por: Zhang, Xiaoyan, et al.
Publicado: (2026)
Ejemplares similares
-
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
por: Liao, Chao, et al.
Publicado: (2025) -
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
por: Li, Kunchang, et al.
Publicado: (2022) -
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
por: Zhou, Chao, et al.
Publicado: (2025) -
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
por: Liu, Jie, et al.
Publicado: (2026) -
VINO: A Unified Visual Generator with Interleaved OmniModal Context
por: Chen, Junyi, et al.
Publicado: (2026)