Data Extrapolation for Text-to-image Generation on Small Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Senmao, Liu, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets
by: Chen, Changjian, et al.
Published: (2024)
by: Chen, Changjian, et al.
Published: (2024)
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026)
by: Liu, Akide, et al.
Published: (2026)
WAS: Dataset and Methods for Artistic Text Segmentation
by: Xie, Xudong, et al.
Published: (2024)
by: Xie, Xudong, et al.
Published: (2024)
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
by: Jelaca, Aleksa, et al.
Published: (2025)
by: Jelaca, Aleksa, et al.
Published: (2025)
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025)
by: Feng, Yukang, et al.
Published: (2025)
Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
by: Kong, Fei
Published: (2025)
by: Kong, Fei
Published: (2025)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
15M Multimodal Facial Image-Text Dataset
by: Dai, Dawei, et al.
Published: (2024)
by: Dai, Dawei, et al.
Published: (2024)
Improving Explicit Spatial Relationships in Text-to-Image Generation through an Automatically Derived Dataset
by: Salaberria, Ander, et al.
Published: (2024)
by: Salaberria, Ander, et al.
Published: (2024)
MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
BlockFusion: Expandable 3D Scene Generation using Latent Tri-plane Extrapolation
by: Wu, Zhennan, et al.
Published: (2024)
by: Wu, Zhennan, et al.
Published: (2024)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Ye, Wei, et al.
Published: (2024)
by: Ye, Wei, et al.
Published: (2024)
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
by: Simonyan, Aleksandr, et al.
Published: (2026)
by: Simonyan, Aleksandr, et al.
Published: (2026)
MUMU: Bootstrapping Multimodal Image Generation from Text-to-Image Data
by: Berman, William, et al.
Published: (2024)
by: Berman, William, et al.
Published: (2024)
From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face Aging
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace
by: Du, Shian, et al.
Published: (2024)
by: Du, Shian, et al.
Published: (2024)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
by: Zhou, Zhiyu, et al.
Published: (2026)
by: Zhou, Zhiyu, et al.
Published: (2026)
When Words Smile: Generating Diverse Emotional Facial Expressions from Text
by: Xu, Haidong, et al.
Published: (2024)
by: Xu, Haidong, et al.
Published: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
by: Ye, Chongjie, et al.
Published: (2026)
by: Ye, Chongjie, et al.
Published: (2026)
LoD Sketch Extraction from Architectural Models Using Generative AI: Dataset Construction for Multi-Level Architectural Design Generation
by: Du, Xusheng, et al.
Published: (2026)
by: Du, Xusheng, et al.
Published: (2026)
CutDiffusion: A Simple, Fast, Cheap, and Strong Diffusion Extrapolation Method
by: Lin, Mingbao, et al.
Published: (2024)
by: Lin, Mingbao, et al.
Published: (2024)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
PointCloud-Text Matching: Benchmark Datasets and a Baseline
by: Feng, Yanglin, et al.
Published: (2024)
by: Feng, Yanglin, et al.
Published: (2024)
HARIVO: Harnessing Text-to-Image Models for Video Generation
by: Kwon, Mingi, et al.
Published: (2024)
by: Kwon, Mingi, et al.
Published: (2024)
RoBus: A Multimodal Dataset for Controllable Road Networks and Building Layouts Generation
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
AccDiffusion v2: Towards More Accurate Higher-Resolution Diffusion Extrapolation
by: Lin, Zhihang, et al.
Published: (2024)
by: Lin, Zhihang, et al.
Published: (2024)
A Generative Approach to High Fidelity 3D Reconstruction from Text Data
by: R, Venkat Kumar, et al.
Published: (2025)
by: R, Venkat Kumar, et al.
Published: (2025)
Multi-language Video Subtitle Dataset for Image-based Text Recognition
by: Singkhornart, Thanadol, et al.
Published: (2024)
by: Singkhornart, Thanadol, et al.
Published: (2024)
Kvasir-VQA: A Text-Image Pair GI Tract Dataset
by: Gautam, Sushant, et al.
Published: (2024)
by: Gautam, Sushant, et al.
Published: (2024)
When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset
by: Das, Sarmistha, et al.
Published: (2025)
by: Das, Sarmistha, et al.
Published: (2025)
FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
by: Cui, Jiacheng, et al.
Published: (2025)
by: Cui, Jiacheng, et al.
Published: (2025)
Addressing Small and Imbalanced Medical Image Datasets Using Generative Models: A Comparative Study of DDPM and PGGANs with Random and Greedy K Sampling
by: Khazrak, Iman, et al.
Published: (2024)
by: Khazrak, Iman, et al.
Published: (2024)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
by: Kou, Siqi, et al.
Published: (2024)
by: Kou, Siqi, et al.
Published: (2024)
Integrated Image-Text Based on Semi-supervised Learning for Small Sample Instance Segmentation
by: Chi, Ruting, et al.
Published: (2024)
by: Chi, Ruting, et al.
Published: (2024)
LACON: Training Text-to-Image Model from Uncurated Data
by: Liang, Zhiyang, et al.
Published: (2026)
by: Liang, Zhiyang, et al.
Published: (2026)
GreenStableYolo: Optimizing Inference Time and Image Quality of Text-to-Image Generation
by: Gong, Jingzhi, et al.
Published: (2024)
by: Gong, Jingzhi, et al.
Published: (2024)
Can Multimodal Large Language Models Truly Understand Small Objects?
by: Han, Fujun, et al.
Published: (2026)
by: Han, Fujun, et al.
Published: (2026)
Similar Items
-
Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets
by: Chen, Changjian, et al.
Published: (2024) -
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026) -
WAS: Dataset and Methods for Artistic Text Segmentation
by: Xie, Xudong, et al.
Published: (2024) -
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
by: Jelaca, Aleksa, et al.
Published: (2025) -
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025)