User-Friendly Customized Generation with Multi-Modal Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhong, Linhao, Hong, Yan, Chen, Wentao, Zhou, Binglin, Zhang, Yiyi, Zhang, Jianfu, Zhang, Liqing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improve Cross-Architecture Generalization on Dataset Distillation
di: Zhou, Binglin, et al.
Pubblicazione: (2024)
di: Zhou, Binglin, et al.
Pubblicazione: (2024)
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
di: Hu, Xinhao, et al.
Pubblicazione: (2026)
di: Hu, Xinhao, et al.
Pubblicazione: (2026)
Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
di: Chen, Tianyi, et al.
Pubblicazione: (2024)
di: Chen, Tianyi, et al.
Pubblicazione: (2024)
High-Quality 3D Head Reconstruction from Any Single Portrait Image
di: Zhang, Jianfu, et al.
Pubblicazione: (2025)
di: Zhang, Jianfu, et al.
Pubblicazione: (2025)
WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection
di: Hong, Yan, et al.
Pubblicazione: (2024)
di: Hong, Yan, et al.
Pubblicazione: (2024)
ComFusion: Personalized Subject Generation in Multiple Specific Scenes From Single Image
di: Hong, Yan, et al.
Pubblicazione: (2024)
di: Hong, Yan, et al.
Pubblicazione: (2024)
Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image
di: Gao, Yujie, et al.
Pubblicazione: (2026)
di: Gao, Yujie, et al.
Pubblicazione: (2026)
DirectTryOn: One-Step Virtual Try-On via Straightened Conditional Transport
di: Sun, Xianbing, et al.
Pubblicazione: (2026)
di: Sun, Xianbing, et al.
Pubblicazione: (2026)
Robustness in AI-Generated Detection: Enhancing Resistance to Adversarial Attacks
di: Haoxuan, Sun, et al.
Pubblicazione: (2025)
di: Haoxuan, Sun, et al.
Pubblicazione: (2025)
DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric Finetuning
di: Duan, Yuxuan, et al.
Pubblicazione: (2024)
di: Duan, Yuxuan, et al.
Pubblicazione: (2024)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
di: Ji, Yikun, et al.
Pubblicazione: (2025)
di: Ji, Yikun, et al.
Pubblicazione: (2025)
DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On
di: Sun, Xianbing, et al.
Pubblicazione: (2025)
di: Sun, Xianbing, et al.
Pubblicazione: (2025)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
di: Ji, Yikun, et al.
Pubblicazione: (2025)
di: Ji, Yikun, et al.
Pubblicazione: (2025)
Double Banking on Knowledge: Customized Modulation and Prototypes for Multi-Modality Semi-supervised Medical Image Segmentation
di: Chen, Yingyu, et al.
Pubblicazione: (2024)
di: Chen, Yingyu, et al.
Pubblicazione: (2024)
WeditGAN: Few-Shot Image Generation via Latent Space Relocation
di: Duan, Yuxuan, et al.
Pubblicazione: (2023)
di: Duan, Yuxuan, et al.
Pubblicazione: (2023)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
di: Hei, Nailei, et al.
Pubblicazione: (2024)
di: Hei, Nailei, et al.
Pubblicazione: (2024)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
di: Chen, Hong, et al.
Pubblicazione: (2024)
di: Chen, Hong, et al.
Pubblicazione: (2024)
Towards Source-Aware Object Swapping with Initial Noise Perturbation
di: Zhan, Jiahui, et al.
Pubblicazione: (2026)
di: Zhan, Jiahui, et al.
Pubblicazione: (2026)
AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality-Missing Prompt Tuning
di: Lang, Jian, et al.
Pubblicazione: (2026)
di: Lang, Jian, et al.
Pubblicazione: (2026)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
di: Chen, Yuheng, et al.
Pubblicazione: (2026)
di: Chen, Yuheng, et al.
Pubblicazione: (2026)
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On
di: Lu, Lingxiao, et al.
Pubblicazione: (2024)
di: Lu, Lingxiao, et al.
Pubblicazione: (2024)
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
di: Wang, Zheng, et al.
Pubblicazione: (2025)
di: Wang, Zheng, et al.
Pubblicazione: (2025)
MSCPT: Few-shot Whole Slide Image Classification with Multi-scale and Context-focused Prompt Tuning
di: Han, Minghao, et al.
Pubblicazione: (2024)
di: Han, Minghao, et al.
Pubblicazione: (2024)
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
di: Wang, Wentao, et al.
Pubblicazione: (2024)
di: Wang, Wentao, et al.
Pubblicazione: (2024)
AVGGT: Rethinking Global Attention for Accelerating VGGT
di: Sun, Xianbing, et al.
Pubblicazione: (2025)
di: Sun, Xianbing, et al.
Pubblicazione: (2025)
Multi-Modal Prompt Learning on Blind Image Quality Assessment
di: Pan, Wensheng, et al.
Pubblicazione: (2024)
di: Pan, Wensheng, et al.
Pubblicazione: (2024)
GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
di: Yan, Haozhen, et al.
Pubblicazione: (2025)
di: Yan, Haozhen, et al.
Pubblicazione: (2025)
Image Captioning via Dynamic Path Customization
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
Multi-Garment Customized Model Generation
di: Liu, Yichen, et al.
Pubblicazione: (2024)
di: Liu, Yichen, et al.
Pubblicazione: (2024)
VTONGuard: Automatic Detection and Authentication of AI-Generated Virtual Try-On Content
di: Wu, Shengyi, et al.
Pubblicazione: (2026)
di: Wu, Shengyi, et al.
Pubblicazione: (2026)
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
di: Kothandaraman, Divya, et al.
Pubblicazione: (2024)
di: Kothandaraman, Divya, et al.
Pubblicazione: (2024)
Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
COMMA: Co-Articulated Multi-Modal Learning
di: Hu, Lianyu, et al.
Pubblicazione: (2023)
di: Hu, Lianyu, et al.
Pubblicazione: (2023)
Multi-Modal Generative Embedding Model
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
di: Guo, Pinxue, et al.
Pubblicazione: (2024)
di: Guo, Pinxue, et al.
Pubblicazione: (2024)
Supervised Contrastive Learning for Snapshot Spectral Imaging Face Anti-Spoofing
di: Song, Chuanbiao, et al.
Pubblicazione: (2024)
di: Song, Chuanbiao, et al.
Pubblicazione: (2024)
Customization Assistant for Text-to-image Generation
di: Zhou, Yufan, et al.
Pubblicazione: (2023)
di: Zhou, Yufan, et al.
Pubblicazione: (2023)
VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models
di: Chen, Hong, et al.
Pubblicazione: (2023)
di: Chen, Hong, et al.
Pubblicazione: (2023)
Pathology-knowledge Enhanced Multi-instance Prompt Learning for Few-shot Whole Slide Image Classification
di: Qu, Linhao, et al.
Pubblicazione: (2024)
di: Qu, Linhao, et al.
Pubblicazione: (2024)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
di: Cao, Guiming, et al.
Pubblicazione: (2024)
di: Cao, Guiming, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Improve Cross-Architecture Generalization on Dataset Distillation
di: Zhou, Binglin, et al.
Pubblicazione: (2024) -
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
di: Hu, Xinhao, et al.
Pubblicazione: (2026) -
Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
di: Chen, Tianyi, et al.
Pubblicazione: (2024) -
High-Quality 3D Head Reconstruction from Any Single Portrait Image
di: Zhang, Jianfu, et al.
Pubblicazione: (2025) -
WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection
di: Hong, Yan, et al.
Pubblicazione: (2024)