Masked Generative Transformer Is What You Need for Image Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Chow, Wei, Li, Linfeng, Sun, Xian, Kong, Lingdong, Li, Zefeng, Xu, Qi, Song, Hang, Ye, Tian, Wang, Xian, Bai, Jinbin, Xu, Shilin, Li, Xiangtai, Pan, Junting, Liu, Shaoteng, Zhou, Ran, Yang, Tianshu, Liu, Songhua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
Not All Points Are Equal: Uncertainty-Aware 4D LiDAR Scene Synthesis
by: Xu, Xiang, et al.
Published: (2026)
by: Xu, Xiang, et al.
Published: (2026)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
by: Xu, Shilin, et al.
Published: (2025)
by: Xu, Shilin, et al.
Published: (2025)
U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
by: Xu, Xiang, et al.
Published: (2025)
by: Xu, Xiang, et al.
Published: (2025)
Not All Tokens Are What You Need In Thinking
by: Yuan, Hang, et al.
Published: (2025)
by: Yuan, Hang, et al.
Published: (2025)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
From Masks to Worlds: A Hitchhiker's Guide to World Models
by: Bai, Jinbin, et al.
Published: (2025)
by: Bai, Jinbin, et al.
Published: (2025)
Landscape of Generative AI in Global News: Topics, Sentiments, and Spatiotemporal Analysis
by: Xian, Lu, et al.
Published: (2024)
by: Xian, Lu, et al.
Published: (2024)
SpotEdit: Selective Region Editing in Diffusion Transformers
by: Qin, Zhibin, et al.
Published: (2025)
by: Qin, Zhibin, et al.
Published: (2025)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Screen‐shooting resistant robust document watermarking in the Discrete Fourier Transform domain
by: Yazhou Zhang, et al.
Published: (2024)
by: Yazhou Zhang, et al.
Published: (2024)
Training on the Benchmark Is Not All You Need
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing
by: Cai, Honghao, et al.
Published: (2026)
by: Cai, Honghao, et al.
Published: (2026)
An Empirical Study of GPT-4o Image Generation Capabilities
by: Chen, Sixiang, et al.
Published: (2025)
by: Chen, Sixiang, et al.
Published: (2025)
FRNet: Frustum-Range Networks for Scalable LiDAR Segmentation
by: Xu, Xiang, et al.
Published: (2023)
by: Xu, Xiang, et al.
Published: (2023)
Edit as You See: Image-guided Video Editing via Masked Motion Modeling
by: Huang, Zhi-Lin, et al.
Published: (2025)
by: Huang, Zhi-Lin, et al.
Published: (2025)
DreamRelation: Bridging Customization and Relation Generation
by: Shi, Qingyu, et al.
Published: (2024)
by: Shi, Qingyu, et al.
Published: (2024)
RecTok: Reconstruction Distillation along Rectified Flow
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
Golden Gemini is All You Need: Finding the Sweet Spots for Speaker Verification
by: Liu, Tianchi, et al.
Published: (2023)
by: Liu, Tianchi, et al.
Published: (2023)
Rho-1: Not All Tokens Are What You Need
by: Lin, Zhenghao, et al.
Published: (2024)
by: Lin, Zhenghao, et al.
Published: (2024)
An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
Is Your LiDAR Placement Optimized for 3D Scene Understanding?
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
by: Li, Jiajun, et al.
Published: (2025)
by: Li, Jiajun, et al.
Published: (2025)
MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale
by: Tang, Zhicong, et al.
Published: (2026)
by: Tang, Zhicong, et al.
Published: (2026)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023)
by: Xu, Shilin, et al.
Published: (2023)
Towards Emergency Scenarios: An Integrated Decision-making Framework of Multi-lane Platoon Reorganization
by: Kong, Aijing, et al.
Published: (2025)
by: Kong, Aijing, et al.
Published: (2025)
Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
by: Li, Leyang, et al.
Published: (2025)
by: Li, Leyang, et al.
Published: (2025)
Monocular Semantic Scene Completion via Masked Recurrent Networks
by: Wang, Xuzhi, et al.
Published: (2025)
by: Wang, Xuzhi, et al.
Published: (2025)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
by: Li, Haodong, et al.
Published: (2026)
by: Li, Haodong, et al.
Published: (2026)
DiTPainter: Efficient Video Inpainting with Diffusion Transformers
by: Wu, Xian, et al.
Published: (2025)
by: Wu, Xian, et al.
Published: (2025)
Veila: Panoramic LiDAR Generation from a Monocular RGB Image
by: Liu, Youquan, et al.
Published: (2025)
by: Liu, Youquan, et al.
Published: (2025)
MaskSR: Masked Language Model for Full-band Speech Restoration
by: Li, Xu, et al.
Published: (2024)
by: Li, Xu, et al.
Published: (2024)
Optimisation Is Not What You Need
by: Ibias, Alfredo
Published: (2025)
by: Ibias, Alfredo
Published: (2025)
Attention Is Not What You Need
by: Chong, Zhang
Published: (2025)
by: Chong, Zhang
Published: (2025)
Similar Items
-
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
by: Chow, Wei, et al.
Published: (2025) -
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
by: Chow, Wei, et al.
Published: (2025) -
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024) -
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
by: Bai, Jinbin, et al.
Published: (2024) -
Not All Points Are Equal: Uncertainty-Aware 4D LiDAR Scene Synthesis
by: Xu, Xiang, et al.
Published: (2026)