Elucidating the design space of language models for image generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xuantong, Hao, Shaozhe, Qi, Xianbiao, Hu, Tianyang, Wang, Jun, Xiao, Rong, Yao, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
by: Hao, Shaozhe, et al.
Published: (2024)
by: Hao, Shaozhe, et al.
Published: (2024)
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
by: Zi, Bojia, et al.
Published: (2025)
by: Zi, Bojia, et al.
Published: (2025)
Refaçade: Editing Object with Given Reference Texture
by: Huang, Youze, et al.
Published: (2025)
by: Huang, Youze, et al.
Published: (2025)
CusConcept: Customized Visual Concept Decomposition with Diffusion Models
by: Xu, Zhi, et al.
Published: (2024)
by: Xu, Zhi, et al.
Published: (2024)
Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
by: He, Wei, et al.
Published: (2026)
by: He, Wei, et al.
Published: (2026)
SimpleGPT: Improving GPT via A Simple Normalization Strategy
by: Chen, Marco, et al.
Published: (2026)
by: Chen, Marco, et al.
Published: (2026)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
by: Wei, Guoting, et al.
Published: (2026)
by: Wei, Guoting, et al.
Published: (2026)
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
by: Zi, Bojia, et al.
Published: (2025)
by: Zi, Bojia, et al.
Published: (2025)
Ctrl&Shift: High-Quality Geometry-Aware Object Manipulation in Visual Generation
by: Ruan, Penghui, et al.
Published: (2026)
by: Ruan, Penghui, et al.
Published: (2026)
Exploring a Principled Framework for Deep Subspace Clustering
by: Meng, Xianghan, et al.
Published: (2025)
by: Meng, Xianghan, et al.
Published: (2025)
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
by: Hao, Shaozhe, et al.
Published: (2024)
by: Hao, Shaozhe, et al.
Published: (2024)
ArtiFade: Learning to Generate High-quality Subject from Blemished Images
by: Yang, Shuya, et al.
Published: (2024)
by: Yang, Shuya, et al.
Published: (2024)
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
CiPR: An Efficient Framework with Cross-instance Positive Relations for Generalized Category Discovery
by: Hao, Shaozhe, et al.
Published: (2023)
by: Hao, Shaozhe, et al.
Published: (2023)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
Double Helix Diffusion for Cross-Domain Anomaly Image Generation
by: Wu, Linchun, et al.
Published: (2025)
by: Wu, Linchun, et al.
Published: (2025)
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
by: Huang, Yaxuan, et al.
Published: (2025)
by: Huang, Yaxuan, et al.
Published: (2025)
Visual Object Tracking on Multi-modal RGB-D Videos: A Review
by: Zhu, Xue-Feng, et al.
Published: (2022)
by: Zhu, Xue-Feng, et al.
Published: (2022)
Winner Team Mia at TextVQA Challenge 2021: Vision-and-Language Representation Learning with Pre-trained Sequence-to-Sequence Model
by: Qiao, Yixuan, et al.
Published: (2021)
by: Qiao, Yixuan, et al.
Published: (2021)
Multi-label Scene Classification for Autonomous Vehicles: Acquiring and Accumulating Knowledge from Diverse Datasets
by: Li, Ke, et al.
Published: (2025)
by: Li, Ke, et al.
Published: (2025)
HMPDM: A Diffusion Model for Driving Video Prediction with Historical Motion Priors
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
Elucidating the solution space of extended reverse-time SDE for diffusion models
by: Cui, Qinpeng, et al.
Published: (2023)
by: Cui, Qinpeng, et al.
Published: (2023)
Taming Transformer Without Using Learning Rate Warmup
by: Qi, Xianbiao, et al.
Published: (2025)
by: Qi, Xianbiao, et al.
Published: (2025)
An Improved Graph Pooling Network for Skeleton-Based Action Recognition
by: Wu, Cong, et al.
Published: (2024)
by: Wu, Cong, et al.
Published: (2024)
Dynamic watermarks in images generated by diffusion models
by: Chen, Yunzhuo, et al.
Published: (2025)
by: Chen, Yunzhuo, et al.
Published: (2025)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)
by: Xu, Lijian, et al.
Published: (2024)
CrashChat: A Multimodal Large Language Model for Multitask Traffic Crash Video Analysis
by: Liang, Kaidi, et al.
Published: (2025)
by: Liang, Kaidi, et al.
Published: (2025)
SalsaAgent: A multimodal embodied language model for interactive dance generation
by: Yazdian, Payam Jome, et al.
Published: (2026)
by: Yazdian, Payam Jome, et al.
Published: (2026)
Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models
by: Li, Yuanbo, et al.
Published: (2026)
by: Li, Yuanbo, et al.
Published: (2026)
Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
by: Li, Yuanbo, et al.
Published: (2026)
by: Li, Yuanbo, et al.
Published: (2026)
Are generative models fair? A study of racial bias in dermatological image generation
by: López-Pérez, Miguel, et al.
Published: (2025)
by: López-Pérez, Miguel, et al.
Published: (2025)
Can video generation replace cinematographers? Research on the cinematic language of generated video
by: Li, Xiaozhe, et al.
Published: (2024)
by: Li, Xiaozhe, et al.
Published: (2024)
Stochastic Interpolants via Conditional Dependent Coupling
by: Ma, Chenrui, et al.
Published: (2025)
by: Ma, Chenrui, et al.
Published: (2025)
IntrinsicEdit: Precise generative image manipulation in intrinsic space
by: Lyu, Linjie, et al.
Published: (2025)
by: Lyu, Linjie, et al.
Published: (2025)
VehicleSDF: A 3D generative model for constrained engineering design via surrogate modeling
by: Morita, Hayata, et al.
Published: (2024)
by: Morita, Hayata, et al.
Published: (2024)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
by: Fu, Yankai, et al.
Published: (2025)
by: Fu, Yankai, et al.
Published: (2025)
A multi-modal vision-language model for generalizable annotation-free pathology localization
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions
by: Ma, Chenrui, et al.
Published: (2025)
by: Ma, Chenrui, et al.
Published: (2025)
Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
by: Zhao, Shihao, et al.
Published: (2025)
by: Zhao, Shihao, et al.
Published: (2025)
Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
by: Huang, Weijian, et al.
Published: (2024)
by: Huang, Weijian, et al.
Published: (2024)
Similar Items
-
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
by: Hao, Shaozhe, et al.
Published: (2024) -
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
by: Zi, Bojia, et al.
Published: (2025) -
Refaçade: Editing Object with Given Reference Texture
by: Huang, Youze, et al.
Published: (2025) -
CusConcept: Customized Visual Concept Decomposition with Diffusion Models
by: Xu, Zhi, et al.
Published: (2024) -
Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
by: He, Wei, et al.
Published: (2026)