Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Image Team, Cai, Huanqia, Cao, Sihan, Du, Ruoyi, Gao, Peng, Hoi, Steven, Hou, Zhaohui, Huang, Shijie, Jiang, Dengyang, Jin, Xin, Li, Liangchen, Li, Zhen, Li, Zhong-Yu, Liu, David, Liu, Dongyang, Shi, Junhan, Wu, Qilong, Yu, Feng, Zhang, Chi, Zhang, Shifeng, Zhou, Shilin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
ImageState collection
by: ImageState
by: ImageState
ImageState collection [Disco compacto] : Browser
by: ImageState
by: ImageState
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
by: Jiang, Dengyang, et al.
Published: (2026)
by: Jiang, Dengyang, et al.
Published: (2026)
Distribution Matching Distillation Meets Reinforcement Learning
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
System-2 Mathematical Reasoning via Enriched Instruction Tuning
by: Cai, Huanqia, et al.
Published: (2024)
by: Cai, Huanqia, et al.
Published: (2024)
VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
by: Li, Zhong-Yu, et al.
Published: (2025)
by: Li, Zhong-Yu, et al.
Published: (2025)
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
by: Li, Haoxin, et al.
Published: (2025)
by: Li, Haoxin, et al.
Published: (2025)
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
by: Huang, Yubo, et al.
Published: (2025)
by: Huang, Yubo, et al.
Published: (2025)
Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval
by: Li, Haiwen, et al.
Published: (2025)
by: Li, Haiwen, et al.
Published: (2025)
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
by: Qin, Qi, et al.
Published: (2025)
by: Qin, Qi, et al.
Published: (2025)
Intelligent Image Search Algorithms Fusing Visual Large Models
by: Wang, Kehan, et al.
Published: (2025)
by: Wang, Kehan, et al.
Published: (2025)
Physics Encoded Spatial and Temporal Generative Adversarial Network for Tropical Cyclone Image Super-resolution
by: Zhang, Ruoyi, et al.
Published: (2026)
by: Zhang, Ruoyi, et al.
Published: (2026)
Layer Construction of Three-Dimensional Z2 Monopole Charge Nodal Line Semimetals and prediction of the abundant candidate materials
by: Li, Yongpan, et al.
Published: (2023)
by: Li, Yongpan, et al.
Published: (2023)
JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Stable Higher-Order Topological Dirac Semimetals with $\mathbb{Z}_2$ Monopole Charge in Alternating-twisted Multilayer Graphenes and beyond
by: Qian, Shifeng, et al.
Published: (2023)
by: Qian, Shifeng, et al.
Published: (2023)
Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
Generative Preprocessing for Image Compression with Pre-trained Diffusion Models
by: Guo, Mengxi, et al.
Published: (2025)
by: Guo, Mengxi, et al.
Published: (2025)
A Robust Multi‐Key Reversible Data Hiding Scheme Leveraging Image Duplication and Smoothness Estimation
by: Zhaohui Li, et al.
Published: (2025)
by: Zhaohui Li, et al.
Published: (2025)
PEARL: Personalized Streaming Video Understanding Model
by: Zheng, Yuanhong, et al.
Published: (2026)
by: Zheng, Yuanhong, et al.
Published: (2026)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow
by: Du, Ruoyi, et al.
Published: (2024)
by: Du, Ruoyi, et al.
Published: (2024)
R3GS: Gaussian Splatting for Robust Reconstruction and Relocalization in Unconstrained Image Collections
by: yan, Xu, et al.
Published: (2025)
by: yan, Xu, et al.
Published: (2025)
Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
A Photolithography‐Free Fabrication Strategy for Perovskite Photodetector Array with High‐Security Imaging Application
by: Jianfeng Zhang, et al.
Published: (2024)
by: Jianfeng Zhang, et al.
Published: (2024)
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
by: Cai, Huanqia, et al.
Published: (2025)
by: Cai, Huanqia, et al.
Published: (2025)
Hierarchical Contrastive Learning for Multimodal Data
by: Li, Huichao, et al.
Published: (2026)
by: Li, Huichao, et al.
Published: (2026)
MARL-MambaContour: Unleashing Multi-Agent Deep Reinforcement Learning for Active Contour Optimization in Medical Image Segmentation
by: Zhang, Ruicheng, et al.
Published: (2025)
by: Zhang, Ruicheng, et al.
Published: (2025)
Enhancing Fundus Image-based Glaucoma Screening via Dynamic Global-Local Feature Integration
by: Zhou, Yuzhuo, et al.
Published: (2025)
by: Zhou, Yuzhuo, et al.
Published: (2025)
Bidirectional Consistency Models
by: Li, Liangchen, et al.
Published: (2024)
by: Li, Liangchen, et al.
Published: (2024)
Baseline Method of the Foundation Model Challenge for Ultrasound Image Analysis
by: Deng, Bo, et al.
Published: (2026)
by: Deng, Bo, et al.
Published: (2026)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
by: Zhao, Shijie, et al.
Published: (2025)
by: Zhao, Shijie, et al.
Published: (2025)
Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Multi-Modal Character Localization and Extraction for Chinese Text Recognition
by: Li, Qilong, et al.
Published: (2026)
by: Li, Qilong, et al.
Published: (2026)
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
ResVR: Joint Rescaling and Viewport Rendering of Omnidirectional Images
by: Li, Weiqi, et al.
Published: (2024)
by: Li, Weiqi, et al.
Published: (2024)
PersonaLive! Expressive Portrait Image Animation for Live Streaming
by: Li, Zhiyuan, et al.
Published: (2025)
by: Li, Zhiyuan, et al.
Published: (2025)
CogStream: Context-guided Streaming Video Question Answering
by: Zhao, Zicheng, et al.
Published: (2025)
by: Zhao, Zicheng, et al.
Published: (2025)
Cover Image, Volume 141, Issue 12
by: Jiuhong Liu, et al.
Published: (2024)
by: Jiuhong Liu, et al.
Published: (2024)
Discrete Element Method–Computational Fluid Dynamics Modeling and Microscopic Failure Mechanism Study of Engineering with Complex Boundary Conditions: A Case Study of a Foundation Pit Project
by: Yuqi Li, et al.
Published: (2025)
by: Yuqi Li, et al.
Published: (2025)
Similar Items
-
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
by: Liu, Dongyang, et al.
Published: (2025) -
ImageState collection
by: ImageState -
ImageState collection [Disco compacto] : Browser
by: ImageState -
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
by: Jiang, Dengyang, et al.
Published: (2026) -
Distribution Matching Distillation Meets Reinforcement Learning
by: Jiang, Dengyang, et al.
Published: (2025)