User-Friendly Customized Generation with Multi-Modal Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Linhao, Hong, Yan, Chen, Wentao, Zhou, Binglin, Zhang, Yiyi, Zhang, Jianfu, Zhang, Liqing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improve Cross-Architecture Generalization on Dataset Distillation
by: Zhou, Binglin, et al.
Published: (2024)
by: Zhou, Binglin, et al.
Published: (2024)
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
by: Hu, Xinhao, et al.
Published: (2026)
by: Hu, Xinhao, et al.
Published: (2026)
Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
by: Chen, Tianyi, et al.
Published: (2024)
by: Chen, Tianyi, et al.
Published: (2024)
High-Quality 3D Head Reconstruction from Any Single Portrait Image
by: Zhang, Jianfu, et al.
Published: (2025)
by: Zhang, Jianfu, et al.
Published: (2025)
WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection
by: Hong, Yan, et al.
Published: (2024)
by: Hong, Yan, et al.
Published: (2024)
ComFusion: Personalized Subject Generation in Multiple Specific Scenes From Single Image
by: Hong, Yan, et al.
Published: (2024)
by: Hong, Yan, et al.
Published: (2024)
Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image
by: Gao, Yujie, et al.
Published: (2026)
by: Gao, Yujie, et al.
Published: (2026)
DirectTryOn: One-Step Virtual Try-On via Straightened Conditional Transport
by: Sun, Xianbing, et al.
Published: (2026)
by: Sun, Xianbing, et al.
Published: (2026)
Robustness in AI-Generated Detection: Enhancing Resistance to Adversarial Attacks
by: Haoxuan, Sun, et al.
Published: (2025)
by: Haoxuan, Sun, et al.
Published: (2025)
DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric Finetuning
by: Duan, Yuxuan, et al.
Published: (2024)
by: Duan, Yuxuan, et al.
Published: (2024)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On
by: Sun, Xianbing, et al.
Published: (2025)
by: Sun, Xianbing, et al.
Published: (2025)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
Double Banking on Knowledge: Customized Modulation and Prototypes for Multi-Modality Semi-supervised Medical Image Segmentation
by: Chen, Yingyu, et al.
Published: (2024)
by: Chen, Yingyu, et al.
Published: (2024)
WeditGAN: Few-Shot Image Generation via Latent Space Relocation
by: Duan, Yuxuan, et al.
Published: (2023)
by: Duan, Yuxuan, et al.
Published: (2023)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
by: Hei, Nailei, et al.
Published: (2024)
by: Hei, Nailei, et al.
Published: (2024)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
by: Chen, Hong, et al.
Published: (2024)
by: Chen, Hong, et al.
Published: (2024)
Towards Source-Aware Object Swapping with Initial Noise Perturbation
by: Zhan, Jiahui, et al.
Published: (2026)
by: Zhan, Jiahui, et al.
Published: (2026)
AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt Tuning
by: Lang, Jian, et al.
Published: (2026)
by: Lang, Jian, et al.
Published: (2026)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On
by: Lu, Lingxiao, et al.
Published: (2024)
by: Lu, Lingxiao, et al.
Published: (2024)
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
by: Wang, Zheng, et al.
Published: (2025)
by: Wang, Zheng, et al.
Published: (2025)
MSCPT: Few-shot Whole Slide Image Classification with Multi-scale and Context-focused Prompt Tuning
by: Han, Minghao, et al.
Published: (2024)
by: Han, Minghao, et al.
Published: (2024)
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
by: Wang, Wentao, et al.
Published: (2024)
by: Wang, Wentao, et al.
Published: (2024)
AVGGT: Rethinking Global Attention for Accelerating VGGT
by: Sun, Xianbing, et al.
Published: (2025)
by: Sun, Xianbing, et al.
Published: (2025)
Multi-Modal Prompt Learning on Blind Image Quality Assessment
by: Pan, Wensheng, et al.
Published: (2024)
by: Pan, Wensheng, et al.
Published: (2024)
GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
by: Yan, Haozhen, et al.
Published: (2025)
by: Yan, Haozhen, et al.
Published: (2025)
Image Captioning via Dynamic Path Customization
by: Ma, Yiwei, et al.
Published: (2024)
by: Ma, Yiwei, et al.
Published: (2024)
Multi-Garment Customized Model Generation
by: Liu, Yichen, et al.
Published: (2024)
by: Liu, Yichen, et al.
Published: (2024)
VTONGuard: Automatic Detection and Authentication of AI-Generated Virtual Try-On Content
by: Wu, Shengyi, et al.
Published: (2026)
by: Wu, Shengyi, et al.
Published: (2026)
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
by: Kothandaraman, Divya, et al.
Published: (2024)
by: Kothandaraman, Divya, et al.
Published: (2024)
Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
COMMA: Co-Articulated Multi-Modal Learning
by: Hu, Lianyu, et al.
Published: (2023)
by: Hu, Lianyu, et al.
Published: (2023)
Multi-Modal Generative Embedding Model
by: Ma, Feipeng, et al.
Published: (2024)
by: Ma, Feipeng, et al.
Published: (2024)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Supervised Contrastive Learning for Snapshot Spectral Imaging Face Anti-Spoofing
by: Song, Chuanbiao, et al.
Published: (2024)
by: Song, Chuanbiao, et al.
Published: (2024)
Customization Assistant for Text-to-image Generation
by: Zhou, Yufan, et al.
Published: (2023)
by: Zhou, Yufan, et al.
Published: (2023)
VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models
by: Chen, Hong, et al.
Published: (2023)
by: Chen, Hong, et al.
Published: (2023)
Pathology-knowledge Enhanced Multi-instance Prompt Learning for Few-shot Whole Slide Image Classification
by: Qu, Linhao, et al.
Published: (2024)
by: Qu, Linhao, et al.
Published: (2024)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
by: Cao, Guiming, et al.
Published: (2024)
by: Cao, Guiming, et al.
Published: (2024)
Similar Items
-
Improve Cross-Architecture Generalization on Dataset Distillation
by: Zhou, Binglin, et al.
Published: (2024) -
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
by: Hu, Xinhao, et al.
Published: (2026) -
Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
by: Chen, Tianyi, et al.
Published: (2024) -
High-Quality 3D Head Reconstruction from Any Single Portrait Image
by: Zhang, Jianfu, et al.
Published: (2025) -
WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection
by: Hong, Yan, et al.
Published: (2024)