CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Cao, Qingqing, Najibi, Mahyar, Mehta, Sachin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
por: Wen, Yuxin, et al.
Publicado: (2024)
por: Wen, Yuxin, et al.
Publicado: (2024)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
por: Mehta, Sachin, et al.
Publicado: (2024)
por: Mehta, Sachin, et al.
Publicado: (2024)
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
por: Huang, Nisha, et al.
Publicado: (2024)
por: Huang, Nisha, et al.
Publicado: (2024)
Ctrl-VI: Controllable Video Synthesis via Variational Inference
por: Duan, Haoyi, et al.
Publicado: (2025)
por: Duan, Haoyi, et al.
Publicado: (2025)
CtrlNeRF: The Generative Neural Radiation Fields for the Controllable Synthesis of High-fidelity 3D-Aware Images
por: Liu, Jian, et al.
Publicado: (2024)
por: Liu, Jian, et al.
Publicado: (2024)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
por: Li, Yaowei, et al.
Publicado: (2025)
por: Li, Yaowei, et al.
Publicado: (2025)
Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes
por: Gosselin, Anthony, et al.
Publicado: (2025)
por: Gosselin, Anthony, et al.
Publicado: (2025)
MUMU: Bootstrapping Multimodal Image Generation from Text-to-Image Data
por: Berman, William, et al.
Publicado: (2024)
por: Berman, William, et al.
Publicado: (2024)
Ctrl-A: Control-Driven Online Data Augmentation
por: Christensen, Jesper B., et al.
Publicado: (2026)
por: Christensen, Jesper B., et al.
Publicado: (2026)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
por: Lin, Han, et al.
Publicado: (2024)
por: Lin, Han, et al.
Publicado: (2024)
DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets
por: Yilmaz, Abdurrahim, et al.
Publicado: (2025)
por: Yilmaz, Abdurrahim, et al.
Publicado: (2025)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
por: Rahman, Kazi Mahathir, et al.
Publicado: (2025)
por: Rahman, Kazi Mahathir, et al.
Publicado: (2025)
Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
por: Chatterjee, Abhiroop, et al.
Publicado: (2025)
por: Chatterjee, Abhiroop, et al.
Publicado: (2025)
Ctrl-GenAug: Controllable Generative Augmentation for Medical Sequence Classification
por: Zhou, Xinrui, et al.
Publicado: (2024)
por: Zhou, Xinrui, et al.
Publicado: (2024)
Efficient Geometry-Controlled High-Resolution Satellite Image Synthesis
por: Vasilescu, Vlad, et al.
Publicado: (2026)
por: Vasilescu, Vlad, et al.
Publicado: (2026)
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
por: Sharifzadeh, Sahand, et al.
Publicado: (2024)
por: Sharifzadeh, Sahand, et al.
Publicado: (2024)
Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
por: Zhang, Letian, et al.
Publicado: (2025)
por: Zhang, Letian, et al.
Publicado: (2025)
Harnessing Shared Relations via Multimodal Mixup Contrastive Learning for Multimodal Classification
por: Kumar, Raja, et al.
Publicado: (2024)
por: Kumar, Raja, et al.
Publicado: (2024)
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization
por: Luo, Yucong, et al.
Publicado: (2024)
por: Luo, Yucong, et al.
Publicado: (2024)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
por: Xu, Shuhan, et al.
Publicado: (2026)
por: Xu, Shuhan, et al.
Publicado: (2026)
Efficient Text-driven Motion Generation via Latent Consistency Training
por: Hu, Mengxian, et al.
Publicado: (2024)
por: Hu, Mengxian, et al.
Publicado: (2024)
KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
por: Lee, Youngwan, et al.
Publicado: (2023)
por: Lee, Youngwan, et al.
Publicado: (2023)
SynthRAR: Ring Artifacts Reduction in CT with Unrolled Network and Synthetic Data Training
por: Yang, Hongxu, et al.
Publicado: (2026)
por: Yang, Hongxu, et al.
Publicado: (2026)
15M Multimodal Facial Image-Text Dataset
por: Dai, Dawei, et al.
Publicado: (2024)
por: Dai, Dawei, et al.
Publicado: (2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023)
por: Wang, Zhouxia, et al.
Publicado: (2023)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
por: Che, Chang, et al.
Publicado: (2024)
por: Che, Chang, et al.
Publicado: (2024)
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
por: Lu, Xiaoxin, et al.
Publicado: (2025)
por: Lu, Xiaoxin, et al.
Publicado: (2025)
Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
por: Joshi, Soham, et al.
Publicado: (2025)
por: Joshi, Soham, et al.
Publicado: (2025)
BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries
por: Li, Tianle, et al.
Publicado: (2025)
por: Li, Tianle, et al.
Publicado: (2025)
Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis
por: Chen, Muxi, et al.
Publicado: (2024)
por: Chen, Muxi, et al.
Publicado: (2024)
Text-only Synthesis for Image Captioning
por: Zhou, Qing, et al.
Publicado: (2024)
por: Zhou, Qing, et al.
Publicado: (2024)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
por: Lu, Guansong, et al.
Publicado: (2023)
por: Lu, Guansong, et al.
Publicado: (2023)
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
por: Dai, Dawei, et al.
Publicado: (2025)
por: Dai, Dawei, et al.
Publicado: (2025)
Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability
por: Zhu, Zhiyu, et al.
Publicado: (2025)
por: Zhu, Zhiyu, et al.
Publicado: (2025)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
por: Peng, Xingkai, et al.
Publicado: (2025)
por: Peng, Xingkai, et al.
Publicado: (2025)
Efficient Egocentric Action Recognition with Multimodal Data
por: Calzavara, Marco, et al.
Publicado: (2025)
por: Calzavara, Marco, et al.
Publicado: (2025)
MULTI: Multimodal Understanding Leaderboard with Text and Images
por: Zhu, Zichen, et al.
Publicado: (2024)
por: Zhu, Zichen, et al.
Publicado: (2024)
TIER: Text-Image Encoder-based Regression for AIGC Image Quality Assessment
por: Yuan, Jiquan, et al.
Publicado: (2024)
por: Yuan, Jiquan, et al.
Publicado: (2024)
EarthSynth: Generating Informative Earth Observation with Diffusion Models
por: Pan, Jiancheng, et al.
Publicado: (2025)
por: Pan, Jiancheng, et al.
Publicado: (2025)
Efficient Exploration of Image Classifier Failures with Bayesian Optimization and Text-to-Image Models
por: LeCoz, Adrien, et al.
Publicado: (2024)
por: LeCoz, Adrien, et al.
Publicado: (2024)
Ejemplares similares
-
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
por: Wen, Yuxin, et al.
Publicado: (2024) -
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
por: Mehta, Sachin, et al.
Publicado: (2024) -
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
por: Huang, Nisha, et al.
Publicado: (2024) -
Ctrl-VI: Controllable Video Synthesis via Variational Inference
por: Duan, Haoyi, et al.
Publicado: (2025) -
CtrlNeRF: The Generative Neural Radiation Fields for the Controllable Synthesis of High-fidelity 3D-Aware Images
por: Liu, Jian, et al.
Publicado: (2024)