Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Shihao, Hao, Shaozhe, Zi, Bojia, Xu, Huaizhe, Wong, Kwan-Yee K. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
di: Hao, Shaozhe, et al.
Pubblicazione: (2024)
di: Hao, Shaozhe, et al.
Pubblicazione: (2024)
Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
di: Zhao, Shihao, et al.
Pubblicazione: (2025)
di: Zhao, Shihao, et al.
Pubblicazione: (2025)
ArtiFade: Learning to Generate High-quality Subject from Blemished Images
di: Yang, Shuya, et al.
Pubblicazione: (2024)
di: Yang, Shuya, et al.
Pubblicazione: (2024)
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
di: Hao, Shaozhe, et al.
Pubblicazione: (2024)
di: Hao, Shaozhe, et al.
Pubblicazione: (2024)
CiPR: An Efficient Framework with Cross-instance Positive Relations for Generalized Category Discovery
di: Hao, Shaozhe, et al.
Pubblicazione: (2023)
di: Hao, Shaozhe, et al.
Pubblicazione: (2023)
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
di: Zi, Bojia, et al.
Pubblicazione: (2025)
di: Zi, Bojia, et al.
Pubblicazione: (2025)
Adversarial Prompt Distillation for Vision-Language Models
di: Luo, Lin, et al.
Pubblicazione: (2024)
di: Luo, Lin, et al.
Pubblicazione: (2024)
A Survey on 3D Human Avatar Modeling -- From Reconstruction to Generation
di: Wang, Ruihe, et al.
Pubblicazione: (2024)
di: Wang, Ruihe, et al.
Pubblicazione: (2024)
CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
di: Zi, Bojia, et al.
Pubblicazione: (2024)
di: Zi, Bojia, et al.
Pubblicazione: (2024)
CusConcept: Customized Visual Concept Decomposition with Diffusion Models
di: Xu, Zhi, et al.
Pubblicazione: (2024)
di: Xu, Zhi, et al.
Pubblicazione: (2024)
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
di: Zi, Bojia, et al.
Pubblicazione: (2025)
di: Zi, Bojia, et al.
Pubblicazione: (2025)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
di: Xie, Chaohao, et al.
Pubblicazione: (2025)
di: Xie, Chaohao, et al.
Pubblicazione: (2025)
InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
VOSR: A Vision-Only Generative Model for Image Super-Resolution
di: Wu, Rongyuan, et al.
Pubblicazione: (2026)
di: Wu, Rongyuan, et al.
Pubblicazione: (2026)
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
di: Lv, Zhengyao, et al.
Pubblicazione: (2025)
di: Lv, Zhengyao, et al.
Pubblicazione: (2025)
CosmicMan: A Text-to-Image Foundation Model for Humans
di: Li, Shikai, et al.
Pubblicazione: (2024)
di: Li, Shikai, et al.
Pubblicazione: (2024)
EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models
di: Zhao, Rui, et al.
Pubblicazione: (2024)
di: Zhao, Rui, et al.
Pubblicazione: (2024)
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
di: Cheng, Hao, et al.
Pubblicazione: (2024)
di: Cheng, Hao, et al.
Pubblicazione: (2024)
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
di: Johnson, Emily, et al.
Pubblicazione: (2025)
di: Johnson, Emily, et al.
Pubblicazione: (2025)
PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis
di: Lv, Zhengyao, et al.
Pubblicazione: (2024)
di: Lv, Zhengyao, et al.
Pubblicazione: (2024)
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
di: Cao, Yukang, et al.
Pubblicazione: (2024)
di: Cao, Yukang, et al.
Pubblicazione: (2024)
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
di: Yang, Hongji, et al.
Pubblicazione: (2025)
di: Yang, Hongji, et al.
Pubblicazione: (2025)
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
MarineEval: Assessing the Marine Intelligence of Vision-Language Models
di: Wong, YuK-Kwan, et al.
Pubblicazione: (2025)
di: Wong, YuK-Kwan, et al.
Pubblicazione: (2025)
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
di: Lan, Mengcheng, et al.
Pubblicazione: (2025)
di: Lan, Mengcheng, et al.
Pubblicazione: (2025)
LooC: Effective Low-Dimensional Codebook for Compositional Vector Quantization
di: Li, Jie, et al.
Pubblicazione: (2026)
di: Li, Jie, et al.
Pubblicazione: (2026)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
di: Li, Yueyan, et al.
Pubblicazione: (2025)
di: Li, Yueyan, et al.
Pubblicazione: (2025)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
di: Fu, Yankai, et al.
Pubblicazione: (2025)
di: Fu, Yankai, et al.
Pubblicazione: (2025)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
di: Lv, Qi, et al.
Pubblicazione: (2025)
di: Lv, Qi, et al.
Pubblicazione: (2025)
Interactive Visual Assessment for Text-to-Image Generation Models
di: Mi, Xiaoyue, et al.
Pubblicazione: (2024)
di: Mi, Xiaoyue, et al.
Pubblicazione: (2024)
Bridging Vision and Language: Optimal Transport-Driven Radiology Report Generation via LLMs
di: Zhao, Haifeng, et al.
Pubblicazione: (2025)
di: Zhao, Haifeng, et al.
Pubblicazione: (2025)
SGEdit: Bridging LLM with Text2Image Generative Model for Scene Graph-based Image Editing
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Zhiyuan, et al.
Pubblicazione: (2024)
Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation
di: Zhang, Xin, et al.
Pubblicazione: (2025)
di: Zhang, Xin, et al.
Pubblicazione: (2025)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
di: Yamabe, Shojiro, et al.
Pubblicazione: (2025)
di: Yamabe, Shojiro, et al.
Pubblicazione: (2025)
An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
TeleStyle: Content-Preserving Style Transfer in Images and Videos
di: Zhang, Shiwen, et al.
Pubblicazione: (2026)
di: Zhang, Shiwen, et al.
Pubblicazione: (2026)
Quantized Prompt for Efficient Generalization of Vision-Language Models
di: Hao, Tianxiang, et al.
Pubblicazione: (2024)
di: Hao, Tianxiang, et al.
Pubblicazione: (2024)
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
di: Basu, Samyadeep, et al.
Pubblicazione: (2024)
di: Basu, Samyadeep, et al.
Pubblicazione: (2024)
Detecting Text Manipulation in Images using Vision Language Models
di: Vidit, Vidit, et al.
Pubblicazione: (2025)
di: Vidit, Vidit, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
di: Hao, Shaozhe, et al.
Pubblicazione: (2024) -
Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
di: Zhao, Shihao, et al.
Pubblicazione: (2025) -
ArtiFade: Learning to Generate High-quality Subject from Blemished Images
di: Yang, Shuya, et al.
Pubblicazione: (2024) -
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
di: Hao, Shaozhe, et al.
Pubblicazione: (2024) -
CiPR: An Efficient Framework with Cross-instance Positive Relations for Generalized Category Discovery
di: Hao, Shaozhe, et al.
Pubblicazione: (2023)