Dynamic Frequency Modulation for Controllable Text-driven Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Tiandong, Zhao, Ling, Qi, Ji, Ma, Jiayi, Peng, Chengli |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
by: Wang, Wenzhuang, et al.
Published: (2025)
by: Wang, Wenzhuang, et al.
Published: (2025)
Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image Models
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model
by: Liu, Chenyang, et al.
Published: (2025)
by: Liu, Chenyang, et al.
Published: (2025)
Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Image Captioning via Dynamic Path Customization
by: Ma, Yiwei, et al.
Published: (2024)
by: Ma, Yiwei, et al.
Published: (2024)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
by: Yang, Danni, et al.
Published: (2024)
by: Yang, Danni, et al.
Published: (2024)
X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation
by: Ma, Yiwei, et al.
Published: (2024)
by: Ma, Yiwei, et al.
Published: (2024)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
by: Yue, Yang, et al.
Published: (2026)
by: Yue, Yang, et al.
Published: (2026)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
by: Yi, Xunpeng, et al.
Published: (2024)
by: Yi, Xunpeng, et al.
Published: (2024)
DyCoRM: Dynamic Criterion-Aware Reward Modeling for Text-to-Image Generation
by: Qian, Jiaying, et al.
Published: (2026)
by: Qian, Jiaying, et al.
Published: (2026)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024)
by: Rahman, Tanzila, et al.
Published: (2024)
X-Dreamer: Creating High-quality 3D Content by Bridging the Domain Gap Between Text-to-2D and Text-to-3D Generation
by: Ma, Yiwei, et al.
Published: (2023)
by: Ma, Yiwei, et al.
Published: (2023)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
by: Wu, Mingrui, et al.
Published: (2026)
by: Wu, Mingrui, et al.
Published: (2026)
Residual Prior-driven Frequency-aware Network for Image Fusion
by: Zheng, Guan, et al.
Published: (2025)
by: Zheng, Guan, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
Exploring Timeline Control for Facial Motion Generation
by: Ma, Yifeng, et al.
Published: (2025)
by: Ma, Yifeng, et al.
Published: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval
by: Ma, Yiwei, et al.
Published: (2024)
by: Ma, Yiwei, et al.
Published: (2024)
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
by: Pang, Lexi, et al.
Published: (2025)
by: Pang, Lexi, et al.
Published: (2025)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
Mixed Degradation Image Restoration via Local Dynamic Optimization and Conditional Embedding
by: Gu, Yubin, et al.
Published: (2024)
by: Gu, Yubin, et al.
Published: (2024)
FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal Model
by: Dang, Shengqi, et al.
Published: (2025)
by: Dang, Shengqi, et al.
Published: (2025)
FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion
by: Yang, Haosen, et al.
Published: (2024)
by: Yang, Haosen, et al.
Published: (2024)
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
by: Lin, Weihuang, et al.
Published: (2025)
by: Lin, Weihuang, et al.
Published: (2025)
Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization
by: Ma, Liyuan, et al.
Published: (2026)
by: Ma, Liyuan, et al.
Published: (2026)
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
by: Pan, Yanjie, et al.
Published: (2025)
by: Pan, Yanjie, et al.
Published: (2025)
Any-to-3D Generation via Hybrid Diffusion Supervision
by: Fan, Yijun, et al.
Published: (2024)
by: Fan, Yijun, et al.
Published: (2024)
TextOVSR: Text-Guided Real-World Opera Video Super-Resolution
by: Chang, Hua, et al.
Published: (2026)
by: Chang, Hua, et al.
Published: (2026)
Text-Animator: Controllable Visual Text Video Generation
by: Liu, Lin, et al.
Published: (2024)
by: Liu, Lin, et al.
Published: (2024)
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
by: Ma, Yiwei, et al.
Published: (2025)
by: Ma, Yiwei, et al.
Published: (2025)
Learning to Sample Effective and Diverse Prompts for Text-to-Image Generation
by: Yun, Taeyoung, et al.
Published: (2025)
by: Yun, Taeyoung, et al.
Published: (2025)
ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation
by: Jin, Hyun-Jun, et al.
Published: (2025)
by: Jin, Hyun-Jun, et al.
Published: (2025)
EvoIR: Towards All-in-One Image Restoration via Evolutionary Frequency Modulation
by: Ma, Jiaqi, et al.
Published: (2025)
by: Ma, Jiaqi, et al.
Published: (2025)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
Similar Items
-
FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
by: Wang, Wenzhuang, et al.
Published: (2025) -
Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
by: Gao, Xiang, et al.
Published: (2024) -
Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image Models
by: Yang, Fan, et al.
Published: (2025) -
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024) -
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)