Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Shilong, Zhang, He, Zhang, Zhifei, Ge, Chongjian, Xue, Shuchen, Liu, Shaoteng, Ren, Mengwei, Kim, Soo Ye, Zhou, Yuqian, Liu, Qing, Pakhomov, Daniil, Zhang, Kai, Lin, Zhe, Luo, Ping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
by: Xue, Shuchen, et al.
Published: (2025)
by: Xue, Shuchen, et al.
Published: (2025)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
Generative Image Layer Decomposition with Visual Effects
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
by: Cai, Yuanhao, et al.
Published: (2024)
by: Cai, Yuanhao, et al.
Published: (2024)
TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting
by: Xie, Liangbin, et al.
Published: (2025)
by: Xie, Liangbin, et al.
Published: (2025)
DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
by: Hong, Susung, et al.
Published: (2025)
by: Hong, Susung, et al.
Published: (2025)
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Rethinking Global Text Conditioning in Diffusion Transformers
by: Starodubcev, Nikita, et al.
Published: (2026)
by: Starodubcev, Nikita, et al.
Published: (2026)
HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
Generative Video Propagation
by: Liu, Shaoteng, et al.
Published: (2024)
by: Liu, Shaoteng, et al.
Published: (2024)
Multitwine: Multi-Object Compositing with Text and Layout Control
by: Tarrés, Gemma Canet, et al.
Published: (2025)
by: Tarrés, Gemma Canet, et al.
Published: (2025)
Object-level Scene Deocclusion
by: Liu, Zhengzhe, et al.
Published: (2024)
by: Liu, Zhengzhe, et al.
Published: (2024)
Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
by: Song, Yizhi, et al.
Published: (2024)
by: Song, Yizhi, et al.
Published: (2024)
TransPixeler: Advancing Text-to-Video Generation with Transparency
by: Wang, Luozhou, et al.
Published: (2025)
by: Wang, Luozhou, et al.
Published: (2025)
Affective Image Editing: Shaping Emotional Factors via Text Descriptions
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
Fusion Makes Perfection: An Efficient Multi-Grained Matching Approach for Zero-Shot Relation Extraction
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
by: Li, Haodong, et al.
Published: (2026)
by: Li, Haodong, et al.
Published: (2026)
Semantic Is Enough: Only Semantic Information For NeRF Reconstruction
by: Wang, Ruibo, et al.
Published: (2024)
by: Wang, Ruibo, et al.
Published: (2024)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
by: Chen, Haoyu, et al.
Published: (2026)
by: Chen, Haoyu, et al.
Published: (2026)
SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing
by: Gu, Jing, et al.
Published: (2024)
by: Gu, Jing, et al.
Published: (2024)
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
by: Zhang, Letian, et al.
Published: (2026)
by: Zhang, Letian, et al.
Published: (2026)
RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing
by: Gong, Yue, et al.
Published: (2026)
by: Gong, Yue, et al.
Published: (2026)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
by: Li, Jilong, et al.
Published: (2024)
by: Li, Jilong, et al.
Published: (2024)
Second-Order Optimality Conditions for Nonsmooth Constrained Optimization with Applications to Bilevel Programming
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Guiding Diffusion Models with Semantically Degraded Conditions
by: Han, Shilong, et al.
Published: (2026)
by: Han, Shilong, et al.
Published: (2026)
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Self-similar blow-up solutions of incompressible Euler equations in $\mathbb R^d, d\geq 3$ with $C^{1,1-2/d-}$ velocity
by: Shao, Feng, et al.
Published: (2026)
by: Shao, Feng, et al.
Published: (2026)
Global derivation of the 1D Vlasov-Poisson equation from quantum many-body dynamics with screened Coulomb potential
by: Chen, Xuwen, et al.
Published: (2024)
by: Chen, Xuwen, et al.
Published: (2024)
Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models
by: Xia, Tao, et al.
Published: (2026)
by: Xia, Tao, et al.
Published: (2026)
The Efficacy of Anterior Cruciate Ligament Reconstruction with Peroneus Longus Tendon and its Impact on Ankle Joint Function
by: Shichao Zhang, et al.
Published: (2024)
by: Shichao Zhang, et al.
Published: (2024)
Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
Text-Guided Semantic Image Encoder
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
SEMDR: A Semantic-Aware Dual Encoder Model for Legal Judgment Prediction with Legal Clue Tracing
by: Liu, Pengjie, et al.
Published: (2024)
by: Liu, Pengjie, et al.
Published: (2024)
Similar Items
-
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
by: Ju, Xuan, et al.
Published: (2025) -
Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
by: Xue, Shuchen, et al.
Published: (2025) -
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025) -
Generative Image Layer Decomposition with Visual Effects
by: Yang, Jinrui, et al.
Published: (2024) -
Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
by: Cai, Yuanhao, et al.
Published: (2024)