Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Jinbin, Ye, Tian, Chow, Wei, Song, Enxin, Li, Xiangtai, Dong, Zhen, Zhu, Lei, Yan, Shuicheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
Masked Generative Transformer Is What You Need for Image Editing
von: Chow, Wei, et al.
Veröffentlicht: (2026)
von: Chow, Wei, et al.
Veröffentlicht: (2026)
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
Integrating View Conditions for Image Synthesis
von: Bai, Jinbin, et al.
Veröffentlicht: (2023)
von: Bai, Jinbin, et al.
Veröffentlicht: (2023)
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer
von: Shao, Shitong, et al.
Veröffentlicht: (2024)
von: Shao, Shitong, et al.
Veröffentlicht: (2024)
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
von: Ye, Tian, et al.
Veröffentlicht: (2025)
von: Ye, Tian, et al.
Veröffentlicht: (2025)
MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer
von: Gao, Shanghua, et al.
Veröffentlicht: (2023)
von: Gao, Shanghua, et al.
Veröffentlicht: (2023)
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
von: Xie, Enze, et al.
Veröffentlicht: (2024)
von: Xie, Enze, et al.
Veröffentlicht: (2024)
From Masks to Worlds: A Hitchhiker's Guide to World Models
von: Bai, Jinbin, et al.
Veröffentlicht: (2025)
von: Bai, Jinbin, et al.
Veröffentlicht: (2025)
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
von: Xu, Weili, et al.
Veröffentlicht: (2025)
von: Xu, Weili, et al.
Veröffentlicht: (2025)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
An Empirical Study of GPT-4o Image Generation Capabilities
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
von: Chen, Sixiang, et al.
Veröffentlicht: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution
von: Liu, Baolin, et al.
Veröffentlicht: (2023)
von: Liu, Baolin, et al.
Veröffentlicht: (2023)
DreamRelation: Bridging Customization and Relation Generation
von: Shi, Qingyu, et al.
Veröffentlicht: (2024)
von: Shi, Qingyu, et al.
Veröffentlicht: (2024)
Personalized Safety Alignment for Text-to-Image Diffusion Models
von: Lei, Yu, et al.
Veröffentlicht: (2025)
von: Lei, Yu, et al.
Veröffentlicht: (2025)
Native-Resolution Image Synthesis
von: Wang, Zidong, et al.
Veröffentlicht: (2025)
von: Wang, Zidong, et al.
Veröffentlicht: (2025)
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution
von: Li, Wenxue, et al.
Veröffentlicht: (2026)
von: Li, Wenxue, et al.
Veröffentlicht: (2026)
DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries
von: Zhou, Yikang, et al.
Veröffentlicht: (2024)
von: Zhou, Yikang, et al.
Veröffentlicht: (2024)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
von: Yu, Chang, et al.
Veröffentlicht: (2024)
von: Yu, Chang, et al.
Veröffentlicht: (2024)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer
von: Fei, Song, et al.
Veröffentlicht: (2025)
von: Fei, Song, et al.
Veröffentlicht: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
Quiz Game Show With PowerPoint and Zoom: Revitalizing a Classical Tool for Modern Learners
von: Nazlee Sharmin, et al.
Veröffentlicht: (2026)
von: Nazlee Sharmin, et al.
Veröffentlicht: (2026)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
DGMamba: Domain Generalization via Generalized State Space Model
von: Long, Shaocong, et al.
Veröffentlicht: (2024)
von: Long, Shaocong, et al.
Veröffentlicht: (2024)
UltraPixel: Advancing Ultra-High-Resolution Image Synthesis to New Peaks
von: Ren, Jingjing, et al.
Veröffentlicht: (2024)
von: Ren, Jingjing, et al.
Veröffentlicht: (2024)
Efficient Transformer for High Resolution Image Motion Deblurring
von: Akmaral, Amanturdieva, et al.
Veröffentlicht: (2025)
von: Akmaral, Amanturdieva, et al.
Veröffentlicht: (2025)
Spikformer V2: Join the High Accuracy Club on ImageNet with an SNN Ticket
von: Zhou, Zhaokun, et al.
Veröffentlicht: (2024)
von: Zhou, Zhaokun, et al.
Veröffentlicht: (2024)
PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2025)
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2025)
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
von: Esser, Patrick, et al.
Veröffentlicht: (2024)
von: Esser, Patrick, et al.
Veröffentlicht: (2024)
InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
von: Han, Tao, et al.
Veröffentlicht: (2025)
von: Han, Tao, et al.
Veröffentlicht: (2025)
PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space Model
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model
von: Huang, Kuan-Chih, et al.
Veröffentlicht: (2024)
von: Huang, Kuan-Chih, et al.
Veröffentlicht: (2024)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2026)
EfficientIML: Efficient High-Resolution Image Manipulation Localization
von: Li, Jinhan, et al.
Veröffentlicht: (2025)
von: Li, Jinhan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
von: Bai, Jinbin, et al.
Veröffentlicht: (2024) -
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
von: Chow, Wei, et al.
Veröffentlicht: (2025) -
Masked Generative Transformer Is What You Need for Image Editing
von: Chow, Wei, et al.
Veröffentlicht: (2026) -
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
von: Shi, Qingyu, et al.
Veröffentlicht: (2025) -
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)