LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Mushui, Ma, Yuhang, Zhen, Yang, Dan, Jun, Yu, Yunlong, Zhao, Zeng, Hu, Zhipeng, Liu, Bai, Fan, Changjie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CM-UNet: Hybrid CNN-Mamba UNet for Remote Sensing Image Semantic Segmentation
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
by: Wu, Fangtai, et al.
Published: (2025)
by: Wu, Fangtai, et al.
Published: (2025)
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
by: He, Weijie, et al.
Published: (2025)
by: He, Weijie, et al.
Published: (2025)
Hybrid Mask Generation for Infrared Small Target Detection with Single-Point Supervision
by: He, Weijie, et al.
Published: (2024)
by: He, Weijie, et al.
Published: (2024)
Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
by: Ma, Yuhang, et al.
Published: (2024)
by: Ma, Yuhang, et al.
Published: (2024)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Fully Fine-tuned CLIP Models are Efficient Few-Shot Learners
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Character-Adapter: Prompt-Guided Region Control for High-Fidelity Character Customization
by: Ma, Yuhang, et al.
Published: (2024)
by: Ma, Yuhang, et al.
Published: (2024)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
by: Li, Bozheng, et al.
Published: (2024)
by: Li, Bozheng, et al.
Published: (2024)
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
by: Ma, Yifeng, et al.
Published: (2023)
by: Ma, Yifeng, et al.
Published: (2023)
Fine-Grained Guidance for Retrievers: Leveraging LLMs' Feedback in Retrieval-Augmented Generation
by: Liu, Yuhang, et al.
Published: (2024)
by: Liu, Yuhang, et al.
Published: (2024)
StdGEN: Semantic-Decomposed 3D Character Generation from Single Images
by: He, Yuze, et al.
Published: (2024)
by: He, Yuze, et al.
Published: (2024)
LTGC: Long-tail Recognition via Leveraging LLMs-driven Generated Content
by: Zhao, Qihao, et al.
Published: (2024)
by: Zhao, Qihao, et al.
Published: (2024)
Towards a Simultaneous and Granular Identity-Expression Control in Personalized Face Generation
by: Liu, Renshuai, et al.
Published: (2024)
by: Liu, Renshuai, et al.
Published: (2024)
Reproducibility Study of "ITI-GEN: Inclusive Text-to-Image Generation"
by: Fernández, Daniel Gallo, et al.
Published: (2024)
by: Fernández, Daniel Gallo, et al.
Published: (2024)
TiP4GEN: Text to Immersive Panorama 4D Scene Generation
by: Xing, Ke, et al.
Published: (2025)
by: Xing, Ke, et al.
Published: (2025)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
by: Liu, Jinlong, et al.
Published: (2026)
by: Liu, Jinlong, et al.
Published: (2026)
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
by: Lin, Hongbin, et al.
Published: (2025)
by: Lin, Hongbin, et al.
Published: (2025)
StdGEN++: A Comprehensive System for Semantic-Decomposed 3D Character Generation
by: He, Yuze, et al.
Published: (2026)
by: He, Yuze, et al.
Published: (2026)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
by: Yi, Xunpeng, et al.
Published: (2024)
by: Yi, Xunpeng, et al.
Published: (2024)
GameUIAgent: An LLM-Powered Framework for Automated Game UI Design with Structured Intermediate Representation
by: Zeng, Wei, et al.
Published: (2026)
by: Zeng, Wei, et al.
Published: (2026)
Envisioning Class Entity Reasoning by Large Language Models for Few-shot Learning
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex Scenes
by: Huang, Feng, et al.
Published: (2025)
by: Huang, Feng, et al.
Published: (2025)
EfficientDreamer: High-Fidelity and Robust 3D Creation via Orthogonal-view Diffusion Prior
by: Hu, Zhipeng, et al.
Published: (2023)
by: Hu, Zhipeng, et al.
Published: (2023)
Learning Interpretable Representations Leads to Semantically Faithful EEG-to-Text Generation
by: Liu, Xiaozhao, et al.
Published: (2025)
by: Liu, Xiaozhao, et al.
Published: (2025)
Game Generation via Large Language Models
by: Hu, Chengpeng, et al.
Published: (2024)
by: Hu, Chengpeng, et al.
Published: (2024)
Dual-Stream Diffusion Net for Text-to-Video Generation
by: Liu, Binhui, et al.
Published: (2023)
by: Liu, Binhui, et al.
Published: (2023)
Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text
by: Najjar, Ayat, et al.
Published: (2025)
by: Najjar, Ayat, et al.
Published: (2025)
Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
by: Liu, Delong, et al.
Published: (2023)
by: Liu, Delong, et al.
Published: (2023)
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
GEN3D: Generating Domain-Free 3D Scenes from a Single Image
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
Compositional Text-to-Image Generation with Dense Blob Representations
by: Nie, Weili, et al.
Published: (2024)
by: Nie, Weili, et al.
Published: (2024)
A General-Purpose Device for Interaction with LLMs
by: Xu, Jiajun, et al.
Published: (2024)
by: Xu, Jiajun, et al.
Published: (2024)
Semantics Disentanglement and Composition for Universal Image Coding with Efficiently LLM Reasoning and Generative Diffusion
by: Liu, Jinming, et al.
Published: (2024)
by: Liu, Jinming, et al.
Published: (2024)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
by: Bai, Jun, et al.
Published: (2025)
by: Bai, Jun, et al.
Published: (2025)
ICE: Interactive 3D Game Character Editing via Dialogue
by: Wu, Haoqian, et al.
Published: (2024)
by: Wu, Haoqian, et al.
Published: (2024)
Annotation-Free Curb Detection Leveraging Altitude Difference Image
by: Ma, Fulong, et al.
Published: (2024)
by: Ma, Fulong, et al.
Published: (2024)
Modification and Generated-Text Detection: Achieving Dual Detection Capabilities for the Outputs of LLM by Watermark
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Similar Items
-
CM-UNet: Hybrid CNN-Mamba UNet for Remote Sensing Image Semantic Segmentation
by: Liu, Mushui, et al.
Published: (2024) -
DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
by: Wu, Fangtai, et al.
Published: (2025) -
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
by: He, Weijie, et al.
Published: (2025) -
Hybrid Mask Generation for Infrared Small Target Detection with Single-Point Supervision
by: He, Weijie, et al.
Published: (2024) -
Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
by: Ma, Yuhang, et al.
Published: (2024)