UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Lunhao, Zhao, Shanshan, Yan, Wenjun, Li, Yinglun, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Gong, Mingming, Xia, Gui-Song |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation
by: Duan, Lunhao, et al.
Published: (2025)
by: Duan, Lunhao, et al.
Published: (2025)
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
by: Chen, Yiyang, et al.
Published: (2025)
by: Chen, Yiyang, et al.
Published: (2025)
Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning
by: Zhou, Sashuai, et al.
Published: (2025)
by: Zhou, Sashuai, et al.
Published: (2025)
Variational Adapter for Cross-modal Similarity Representation
by: Wei, WenZhang, et al.
Published: (2026)
by: Wei, WenZhang, et al.
Published: (2026)
StyleAdapter: A Unified Stylized Image Generation Model
by: Wang, Zhouxia, et al.
Published: (2023)
by: Wang, Zhouxia, et al.
Published: (2023)
Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis
by: Chen, Yiyang, et al.
Published: (2024)
by: Chen, Yiyang, et al.
Published: (2024)
UNIC: Unified In-Context Video Editing
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
by: Song, Kunpeng, et al.
Published: (2024)
by: Song, Kunpeng, et al.
Published: (2024)
Federated Client-tailored Adapter for Medical Image Segmentation
by: Hu, Guyue, et al.
Published: (2025)
by: Hu, Guyue, et al.
Published: (2025)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
HAMUR: Hyper Adapter for Multi-Domain Recommendation
by: Li, Xiaopeng, et al.
Published: (2023)
by: Li, Xiaopeng, et al.
Published: (2023)
MV-Adapter: Multi-view Consistent Image Generation Made Easy
by: Huang, Zehuan, et al.
Published: (2024)
by: Huang, Zehuan, et al.
Published: (2024)
Adapter-dependent Adapter Methylation Assay
by: Zhang, Jia, et al.
Published: (2024)
by: Zhang, Jia, et al.
Published: (2024)
ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
by: Dong, Sibo, et al.
Published: (2025)
by: Dong, Sibo, et al.
Published: (2025)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
by: Guo, Xun, et al.
Published: (2023)
by: Guo, Xun, et al.
Published: (2023)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Adapter
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision
by: Chen, Jiacheng, et al.
Published: (2026)
by: Chen, Jiacheng, et al.
Published: (2026)
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
by: Wang, Yiqun, et al.
Published: (2024)
by: Wang, Yiqun, et al.
Published: (2024)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
by: Zhao, Yumiao, et al.
Published: (2024)
by: Zhao, Yumiao, et al.
Published: (2024)
CAT: Contrastive Adapter Training for Personalized Image Generation
by: Park, Jae Wan, et al.
Published: (2024)
by: Park, Jae Wan, et al.
Published: (2024)
On the Duality between Gradient Transformations and Adapters
by: Torroba-Hennigen, Lucas, et al.
Published: (2025)
by: Torroba-Hennigen, Lucas, et al.
Published: (2025)
Text to Image for Multi-Label Image Recognition with Joint Prompt-Adapter Learning
by: Feng, Chun-Mei, et al.
Published: (2025)
by: Feng, Chun-Mei, et al.
Published: (2025)
Memory-based Adapters for Online 3D Scene Perception
by: Xu, Xiuwei, et al.
Published: (2024)
by: Xu, Xiuwei, et al.
Published: (2024)
LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Image
by: Madan, Chetan, et al.
Published: (2024)
by: Madan, Chetan, et al.
Published: (2024)
UNIC: Learning Unified Multimodal Extrinsic Contact Estimation
by: Xu, Zhengtong, et al.
Published: (2026)
by: Xu, Zhengtong, et al.
Published: (2026)
IDEA: Image Description Enhanced CLIP-Adapter
by: Ye, Zhipeng, et al.
Published: (2025)
by: Ye, Zhipeng, et al.
Published: (2025)
Cross-modality Attention Adapter: A Glioma Segmentation Fine-tuning Method for SAM Using Multimodal Brain MR Images
by: Shi, Xiaoyu, et al.
Published: (2023)
by: Shi, Xiaoyu, et al.
Published: (2023)
Memory Efficient Transformer Adapter for Dense Predictions
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
AdapterTune: Zero-Initialized Low-Rank Adapters for Frozen Vision Transformers
by: Khazem, Salim
Published: (2026)
by: Khazem, Salim
Published: (2026)
DP-Adapter: Dual-Pathway Adapter for Boosting Fidelity and Text Consistency in Customizable Human Image Generation
by: Wang, Ye, et al.
Published: (2025)
by: Wang, Ye, et al.
Published: (2025)
ResAdapter: Domain Consistent Resolution Adapter for Diffusion Models
by: Cheng, Jiaxiang, et al.
Published: (2024)
by: Cheng, Jiaxiang, et al.
Published: (2024)
Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation
by: Jia, Jun, et al.
Published: (2025)
by: Jia, Jun, et al.
Published: (2025)
SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Frequency Adapter with SAM for Generalized Medical Image Segmentation
by: Bui, Phuoc-Nguyen, et al.
Published: (2026)
by: Bui, Phuoc-Nguyen, et al.
Published: (2026)
Similar Items
-
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025) -
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025) -
High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation
by: Duan, Lunhao, et al.
Published: (2025) -
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
by: Chen, Yiyang, et al.
Published: (2025) -
Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning
by: Zhou, Sashuai, et al.
Published: (2025)