UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Tian, Fei, Song, Zhu, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer
by: Fei, Song, et al.
Published: (2025)
by: Fei, Song, et al.
Published: (2025)
NARAIM: Native Aspect Ratio Autoregressive Image Models
by: Fernández, Daniel Gallo, et al.
Published: (2024)
by: Fernández, Daniel Gallo, et al.
Published: (2024)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
by: Gao, Zhanxin, et al.
Published: (2025)
by: Gao, Zhanxin, et al.
Published: (2025)
Data Extrapolation for Text-to-image Generation on Small Datasets
by: Ye, Senmao, et al.
Published: (2024)
by: Ye, Senmao, et al.
Published: (2024)
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution
by: Li, Wenxue, et al.
Published: (2026)
by: Li, Wenxue, et al.
Published: (2026)
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
by: Xu, Ruyi, et al.
Published: (2024)
by: Xu, Ruyi, et al.
Published: (2024)
Diverse Text-to-Image Generation via Contrastive Noise Optimization
by: Kim, Byungjun, et al.
Published: (2025)
by: Kim, Byungjun, et al.
Published: (2025)
UltraPixel: Advancing Ultra-High-Resolution Image Synthesis to New Peaks
by: Ren, Jingjing, et al.
Published: (2024)
by: Ren, Jingjing, et al.
Published: (2024)
CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities
by: Gong, Yue, et al.
Published: (2025)
by: Gong, Yue, et al.
Published: (2025)
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
by: Wang, Zhendong, et al.
Published: (2025)
by: Wang, Zhendong, et al.
Published: (2025)
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
by: Zhang, Jinjin, et al.
Published: (2025)
by: Zhang, Jinjin, et al.
Published: (2025)
DescriptorMedSAM: Language-Image Fusion with Multi-Aspect Text Guidance for Medical Image Segmentation
by: Zhang, Wenjie, et al.
Published: (2025)
by: Zhang, Wenjie, et al.
Published: (2025)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
by: Ba, Ying, et al.
Published: (2025)
by: Ba, Ying, et al.
Published: (2025)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
Learning to Sample Effective and Diverse Prompts for Text-to-Image Generation
by: Yun, Taeyoung, et al.
Published: (2025)
by: Yun, Taeyoung, et al.
Published: (2025)
D2D: Detector-to-Differentiable Critic for Improved Numeracy in Text-to-Image Generation
by: Yoo, Nobline, et al.
Published: (2025)
by: Yoo, Nobline, et al.
Published: (2025)
Preliminary Explorations with GPT-4o(mni) Native Image Generation
by: Cao, Pu, et al.
Published: (2025)
by: Cao, Pu, et al.
Published: (2025)
Generating Multi-Image Synthetic Data for Text-to-Image Customization
by: Kumari, Nupur, et al.
Published: (2025)
by: Kumari, Nupur, et al.
Published: (2025)
LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution
by: Fei, Song, et al.
Published: (2026)
by: Fei, Song, et al.
Published: (2026)
Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
by: Ren, Jingjing, et al.
Published: (2025)
by: Ren, Jingjing, et al.
Published: (2025)
Navigating Text-to-Image Generative Bias across Indic Languages
by: Mittal, Surbhi, et al.
Published: (2024)
by: Mittal, Surbhi, et al.
Published: (2024)
Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation
by: Wu, Qingxuan, et al.
Published: (2025)
by: Wu, Qingxuan, et al.
Published: (2025)
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model
by: Gong, Lixue, et al.
Published: (2025)
by: Gong, Lixue, et al.
Published: (2025)
Native-Resolution Image Synthesis
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
DynamiCtrl: Rethinking the Basic Structure and the Role of Text for High-quality Human Image Animation
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment
by: Gao, Zhiting, et al.
Published: (2025)
by: Gao, Zhiting, et al.
Published: (2025)
PKU-AIGIQA-4K: A Perceptual Quality Assessment Database for Both Text-to-Image and Image-to-Image AI-Generated Images
by: Yuan, Jiquan, et al.
Published: (2024)
by: Yuan, Jiquan, et al.
Published: (2024)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
by: Chern, Ethan, et al.
Published: (2024)
by: Chern, Ethan, et al.
Published: (2024)
Measuring Diversity in Co-creative Image Generation
by: Ibarrola, Francisco, et al.
Published: (2024)
by: Ibarrola, Francisco, et al.
Published: (2024)
QuadGPT: Native Quadrilateral Mesh Generation with Autoregressive Models
by: Liu, Jian, et al.
Published: (2025)
by: Liu, Jian, et al.
Published: (2025)
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
by: Xie, Yu, et al.
Published: (2025)
by: Xie, Yu, et al.
Published: (2025)
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
by: Zhao, Chenxi, et al.
Published: (2026)
by: Zhao, Chenxi, et al.
Published: (2026)
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
by: Wang, Weizhi, et al.
Published: (2024)
by: Wang, Weizhi, et al.
Published: (2024)
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
Text4Seg: Reimagining Image Segmentation as Text Generation
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
Generalize Polyp Segmentation via Inpainting across Diverse Backgrounds and Pseudo-Mask Refinement
by: Ma, Jiajian, et al.
Published: (2024)
by: Ma, Jiajian, et al.
Published: (2024)
Similar Items
-
LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer
by: Fei, Song, et al.
Published: (2025) -
NARAIM: Native Aspect Ratio Autoregressive Image Models
by: Fernández, Daniel Gallo, et al.
Published: (2024) -
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024) -
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
by: Gao, Zhanxin, et al.
Published: (2025) -
Data Extrapolation for Text-to-image Generation on Small Datasets
by: Ye, Senmao, et al.
Published: (2024)