ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Sibo, Shaheen, Ismail, Shen, Maggie, Mallick, Rupayan, Bargal, Sarah Adel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
D-Feat Occlusions: Diffusion Features for Robustness to Partial Visual Occlusions in Object Recognition
by: Mallick, Rupayan, et al.
Published: (2025)
by: Mallick, Rupayan, et al.
Published: (2025)
FaithFill: Faithful Inpainting for Object Completion Using a Single Reference Image
by: Mallick, Rupayan, et al.
Published: (2024)
by: Mallick, Rupayan, et al.
Published: (2024)
GenEAva: Generating Cartoon Avatars with Fine-Grained Facial Expressions from Realistic Diffusion-based Faces
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
by: Zhang, Ganlin, et al.
Published: (2025)
by: Zhang, Ganlin, et al.
Published: (2025)
GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention
by: Abdalla, Amro, et al.
Published: (2025)
by: Abdalla, Amro, et al.
Published: (2025)
Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
by: Duan, Lunhao, et al.
Published: (2024)
by: Duan, Lunhao, et al.
Published: (2024)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
by: Wu, Linquan, et al.
Published: (2026)
by: Wu, Linquan, et al.
Published: (2026)
A Manually Annotated Image-Caption Dataset for Detecting Children in the Wild
by: Kireev, Klim, et al.
Published: (2025)
by: Kireev, Klim, et al.
Published: (2025)
LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Image
by: Madan, Chetan, et al.
Published: (2024)
by: Madan, Chetan, et al.
Published: (2024)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
by: Liu, Jiaming, et al.
Published: (2023)
by: Liu, Jiaming, et al.
Published: (2023)
Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization
by: Zhang, Yuxi, et al.
Published: (2025)
by: Zhang, Yuxi, et al.
Published: (2025)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
by: Li, Zhiyang, et al.
Published: (2026)
by: Li, Zhiyang, et al.
Published: (2026)
Text to Image for Multi-Label Image Recognition with Joint Prompt-Adapter Learning
by: Feng, Chun-Mei, et al.
Published: (2025)
by: Feng, Chun-Mei, et al.
Published: (2025)
Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
From Image Captioning to Visual Storytelling
by: Passadakis, Admitos, et al.
Published: (2025)
by: Passadakis, Admitos, et al.
Published: (2025)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
by: Jiang, Yue, et al.
Published: (2026)
by: Jiang, Yue, et al.
Published: (2026)
Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual Generation
by: Tong, Lei, et al.
Published: (2025)
by: Tong, Lei, et al.
Published: (2025)
ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit Adaptation
by: Chen, Dar-Yen, et al.
Published: (2023)
by: Chen, Dar-Yen, et al.
Published: (2023)
Efficient Text-Guided Convolutional Adapter for the Diffusion Model
by: Das, Aryan, et al.
Published: (2026)
by: Das, Aryan, et al.
Published: (2026)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
by: Jha, Akshita, et al.
Published: (2024)
by: Jha, Akshita, et al.
Published: (2024)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval
by: Cai, Rui, et al.
Published: (2024)
by: Cai, Rui, et al.
Published: (2024)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
by: Shen, Guibao, et al.
Published: (2024)
by: Shen, Guibao, et al.
Published: (2024)
FontAdapter: Instant Font Adaptation in Visual Text Generation
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
CoSTA$\ast$: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing
by: Gupta, Advait, et al.
Published: (2025)
by: Gupta, Advait, et al.
Published: (2025)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
by: Guo, Xun, et al.
Published: (2023)
by: Guo, Xun, et al.
Published: (2023)
An Intermediate Fusion ViT Enables Efficient Text-Image Alignment in Diffusion Models
by: Hu, Zizhao, et al.
Published: (2024)
by: Hu, Zizhao, et al.
Published: (2024)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models
by: Sreenivas, Manogna, et al.
Published: (2026)
by: Sreenivas, Manogna, et al.
Published: (2026)
Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models
by: Shen, Fei, et al.
Published: (2024)
by: Shen, Fei, et al.
Published: (2024)
Diffusion Restoration Adapter for Real-World Image Restoration
by: Liang, Hanbang, et al.
Published: (2025)
by: Liang, Hanbang, et al.
Published: (2025)
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in non-English Text-to-Image Generation
by: Ma, Jian, et al.
Published: (2023)
by: Ma, Jian, et al.
Published: (2023)
Similar Items
-
D-Feat Occlusions: Diffusion Features for Robustness to Partial Visual Occlusions in Object Recognition
by: Mallick, Rupayan, et al.
Published: (2025) -
FaithFill: Faithful Inpainting for Object Completion Using a Single Reference Image
by: Mallick, Rupayan, et al.
Published: (2024) -
GenEAva: Generating Cartoon Avatars with Fine-Grained Facial Expressions from Realistic Diffusion-based Faces
by: Yu, Hao, et al.
Published: (2025) -
ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
by: Zhang, Ganlin, et al.
Published: (2025) -
GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention
by: Abdalla, Amro, et al.
Published: (2025)