CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Nan, Huang, Mengqi, Chen, Zhuowei, Zheng, Yang, Zhang, Lei, Mao, Zhendong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
di: Huang, Mengqi, et al.
Pubblicazione: (2024)
di: Huang, Mengqi, et al.
Pubblicazione: (2024)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
di: Chen, Yuheng, et al.
Pubblicazione: (2026)
di: Chen, Yuheng, et al.
Pubblicazione: (2026)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
di: Mao, Zhendong, et al.
Pubblicazione: (2024)
di: Mao, Zhendong, et al.
Pubblicazione: (2024)
Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition
di: Chang, Haochen, et al.
Pubblicazione: (2024)
di: Chang, Haochen, et al.
Pubblicazione: (2024)
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
di: Lin, Honglin, et al.
Pubblicazione: (2024)
di: Lin, Honglin, et al.
Pubblicazione: (2024)
PuLID: Pure and Lightning ID Customization via Contrastive Alignment
di: Guo, Zinan, et al.
Pubblicazione: (2024)
di: Guo, Zinan, et al.
Pubblicazione: (2024)
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
di: Shan, Ziyu, et al.
Pubblicazione: (2024)
di: Shan, Ziyu, et al.
Pubblicazione: (2024)
Exploring Rich Subjective Quality Information for Image Quality Assessment in the Wild
di: Min, Xiongkuo, et al.
Pubblicazione: (2024)
di: Min, Xiongkuo, et al.
Pubblicazione: (2024)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
di: Xu, Yu, et al.
Pubblicazione: (2024)
di: Xu, Yu, et al.
Pubblicazione: (2024)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
di: Yang, Shuyu, et al.
Pubblicazione: (2024)
di: Yang, Shuyu, et al.
Pubblicazione: (2024)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
di: Zhang, Rui, et al.
Pubblicazione: (2024)
di: Zhang, Rui, et al.
Pubblicazione: (2024)
G-Refine: A General Quality Refiner for Text-to-Image Generation
di: Li, Chunyi, et al.
Pubblicazione: (2024)
di: Li, Chunyi, et al.
Pubblicazione: (2024)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
di: Xiao, Jian, et al.
Pubblicazione: (2025)
di: Xiao, Jian, et al.
Pubblicazione: (2025)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
di: Qin, Yang, et al.
Pubblicazione: (2023)
di: Qin, Yang, et al.
Pubblicazione: (2023)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
di: Chen, Junyu, et al.
Pubblicazione: (2025)
di: Chen, Junyu, et al.
Pubblicazione: (2025)
Deep Contrastive Multi-view Clustering under Semantic Feature Guidance
di: Liu, Siwen, et al.
Pubblicazione: (2024)
di: Liu, Siwen, et al.
Pubblicazione: (2024)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
di: Gao, Jiayi, et al.
Pubblicazione: (2025)
di: Gao, Jiayi, et al.
Pubblicazione: (2025)
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
di: Baraldi, Lorenzo, et al.
Pubblicazione: (2024)
di: Baraldi, Lorenzo, et al.
Pubblicazione: (2024)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
di: Yang, Jianxuan, et al.
Pubblicazione: (2025)
di: Yang, Jianxuan, et al.
Pubblicazione: (2025)
A Simple Task-aware Contrastive Local Descriptor Selection Strategy for Few-shot Learning between inter class and intra class
di: Qiao, Qian, et al.
Pubblicazione: (2024)
di: Qiao, Qian, et al.
Pubblicazione: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
di: Jiang, Chen, et al.
Pubblicazione: (2023)
di: Jiang, Chen, et al.
Pubblicazione: (2023)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
di: Chen, Liyang, et al.
Pubblicazione: (2025)
di: Chen, Liyang, et al.
Pubblicazione: (2025)
Regularized Contrastive Partial Multi-view Outlier Detection
di: Wang, Yijia, et al.
Pubblicazione: (2024)
di: Wang, Yijia, et al.
Pubblicazione: (2024)
Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
di: Fang, Ziye, et al.
Pubblicazione: (2023)
di: Fang, Ziye, et al.
Pubblicazione: (2023)
Med-Banana-50K: A Cross-modality Large-Scale Dataset for Text-guided Medical Image Editing
di: Chen, Zhihui, et al.
Pubblicazione: (2025)
di: Chen, Zhihui, et al.
Pubblicazione: (2025)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
di: Huang, Yuesheng, et al.
Pubblicazione: (2025)
di: Huang, Yuesheng, et al.
Pubblicazione: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
di: Xu, Jingning, et al.
Pubblicazione: (2026)
di: Xu, Jingning, et al.
Pubblicazione: (2026)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
di: Deng, Jingyuan, et al.
Pubblicazione: (2025)
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
di: Gong, Zixuan, et al.
Pubblicazione: (2024)
di: Gong, Zixuan, et al.
Pubblicazione: (2024)
Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
di: Yin, Qilin, et al.
Pubblicazione: (2025)
di: Yin, Qilin, et al.
Pubblicazione: (2025)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026)
di: Song, Shezheng, et al.
Pubblicazione: (2026)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
di: Huang, Hailang, et al.
Pubblicazione: (2024)
di: Huang, Hailang, et al.
Pubblicazione: (2024)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
di: Wu, Xun, et al.
Pubblicazione: (2024)
di: Wu, Xun, et al.
Pubblicazione: (2024)
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
di: Le, Anh-Duy, et al.
Pubblicazione: (2026)
di: Le, Anh-Duy, et al.
Pubblicazione: (2026)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
di: Swetha, Sirnam, et al.
Pubblicazione: (2024)
AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation
di: Chen, Liyang, et al.
Pubblicazione: (2023)
di: Chen, Liyang, et al.
Pubblicazione: (2023)
Magic3DSketch: Create Colorful 3D Models From Sketch-Based 3D Modeling Guided by Text and Language-Image Pre-Training
di: Zang, Ying, et al.
Pubblicazione: (2024)
di: Zang, Ying, et al.
Pubblicazione: (2024)
Documenti analoghi
-
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
di: Huang, Mengqi, et al.
Pubblicazione: (2024) -
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
di: Chen, Yuheng, et al.
Pubblicazione: (2026) -
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025) -
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025) -
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
di: Mao, Zhendong, et al.
Pubblicazione: (2024)