Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Chenxi, Zhu, Chen, Feng, Xiaokun, Hao, Aiming, Zhu, Jiashu, Lei, Jiachen, Wu, Jiahong, Chu, Xiangxiang, Yang, Jufeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ranking-aware Reinforcement Learning for Ordinal Ranking
von: Hao, Aiming, et al.
Veröffentlicht: (2026)
von: Hao, Aiming, et al.
Veröffentlicht: (2026)
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
von: Ling, Xinran, et al.
Veröffentlicht: (2025)
von: Ling, Xinran, et al.
Veröffentlicht: (2025)
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
von: Mao, Fangyuan, et al.
Veröffentlicht: (2025)
von: Mao, Fangyuan, et al.
Veröffentlicht: (2025)
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2026)
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
Stochastic Self-Guidance for Training-Free Enhancement of Diffusion Models
von: Chen, Chubin, et al.
Veröffentlicht: (2025)
von: Chen, Chubin, et al.
Veröffentlicht: (2025)
Embedding-perturbed Exploration Preference Optimization for Flow Models
von: Hu, Sujie, et al.
Veröffentlicht: (2026)
von: Hu, Sujie, et al.
Veröffentlicht: (2026)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
One-Step Image Translation with Text-to-Image Models
von: Parmar, Gaurav, et al.
Veröffentlicht: (2024)
von: Parmar, Gaurav, et al.
Veröffentlicht: (2024)
Interpretable Discriminative Text Representations via Agreement and Label Disentanglement
von: Wang, Tong, et al.
Veröffentlicht: (2026)
von: Wang, Tong, et al.
Veröffentlicht: (2026)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
von: Lei, Jiachen, et al.
Veröffentlicht: (2025)
von: Lei, Jiachen, et al.
Veröffentlicht: (2025)
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
Artifact-Aware Evaluation for High-Quality Video Generation
von: Zhu, Chen, et al.
Veröffentlicht: (2026)
von: Zhu, Chen, et al.
Veröffentlicht: (2026)
Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
von: Dao, Quan, et al.
Veröffentlicht: (2024)
von: Dao, Quan, et al.
Veröffentlicht: (2024)
Discriminative Class Tokens for Text-to-Image Diffusion Models
von: Schwartz, Idan, et al.
Veröffentlicht: (2023)
von: Schwartz, Idan, et al.
Veröffentlicht: (2023)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
FlowDreamer: Exploring High Fidelity Text-to-3D Generation via Rectified Flow
von: Li, Hangyu, et al.
Veröffentlicht: (2024)
von: Li, Hangyu, et al.
Veröffentlicht: (2024)
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models
von: Wang, Zengbin, et al.
Veröffentlicht: (2026)
von: Wang, Zengbin, et al.
Veröffentlicht: (2026)
ConceptWeaver: Weaving Disentangled Concepts with Flow
von: Chen, Jintao, et al.
Veröffentlicht: (2026)
von: Chen, Jintao, et al.
Veröffentlicht: (2026)
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
von: Yang, Kaixing, et al.
Veröffentlicht: (2025)
Discriminative Probing and Tuning for Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution
von: He, Haodong, et al.
Veröffentlicht: (2026)
von: He, Haodong, et al.
Veröffentlicht: (2026)
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2026)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2026)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
von: Ma, Leilei, et al.
Veröffentlicht: (2024)
von: Ma, Leilei, et al.
Veröffentlicht: (2024)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning
von: Chen, Chubin, et al.
Veröffentlicht: (2025)
von: Chen, Chubin, et al.
Veröffentlicht: (2025)
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark
von: Jiang, Feng, et al.
Veröffentlicht: (2023)
von: Jiang, Feng, et al.
Veröffentlicht: (2023)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
von: Yu, Hu, et al.
Veröffentlicht: (2024)
von: Yu, Hu, et al.
Veröffentlicht: (2024)
Expressive Text-to-Image Generation with Rich Text
von: Ge, Songwei, et al.
Veröffentlicht: (2023)
von: Ge, Songwei, et al.
Veröffentlicht: (2023)
Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2024)
Leave-One-Out Prediction for General Hypothesis Classes
von: Qian, Jian, et al.
Veröffentlicht: (2026)
von: Qian, Jian, et al.
Veröffentlicht: (2026)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness
von: Chen, Honghao, et al.
Veröffentlicht: (2024)
von: Chen, Honghao, et al.
Veröffentlicht: (2024)
Representations of Text and Images Align From Layer One
von: Wybitul, Evžen, et al.
Veröffentlicht: (2026)
von: Wybitul, Evžen, et al.
Veröffentlicht: (2026)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
von: Surkov, Viacheslav, et al.
Veröffentlicht: (2024)
von: Surkov, Viacheslav, et al.
Veröffentlicht: (2024)
InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation
von: Liu, Xingchao, et al.
Veröffentlicht: (2023)
von: Liu, Xingchao, et al.
Veröffentlicht: (2023)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Tung, et al.
Veröffentlicht: (2024)
WIFE-Fusion:Wavelet-aware Intra-inter Frequency Enhancement for Multi-model Image Fusion
von: Zhang, Tianpei, et al.
Veröffentlicht: (2025)
von: Zhang, Tianpei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ranking-aware Reinforcement Learning for Ordinal Ranking
von: Hao, Aiming, et al.
Veröffentlicht: (2026) -
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
von: Ling, Xinran, et al.
Veröffentlicht: (2025) -
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
von: Mao, Fangyuan, et al.
Veröffentlicht: (2025) -
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2026) -
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)