Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yufan, Zhang, Ruiyi, Zheng, Kaizhi, Zhao, Nanxuan, Gu, Jiuxiang, Wang, Zichao, Wang, Xin Eric, Sun, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner
von: Zhou, Yufan, et al.
Veröffentlicht: (2024)
von: Zhou, Yufan, et al.
Veröffentlicht: (2024)
Customization Assistant for Text-to-image Generation
von: Zhou, Yufan, et al.
Veröffentlicht: (2023)
von: Zhou, Yufan, et al.
Veröffentlicht: (2023)
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
von: Zhou, Shijie, et al.
Veröffentlicht: (2024)
TextLap: Customizing Language Models for Text-to-Layout Planning
von: Chen, Jian, et al.
Veröffentlicht: (2024)
von: Chen, Jian, et al.
Veröffentlicht: (2024)
Towards Visual Text Grounding of Multimodal Large Language Model
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
TRINS: Towards Multimodal Language Models that Can Read
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
Constructing a 3D Scene from a Single Image
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
MMR: Evaluating Reading Ability of Large Multimodal Models
von: Chen, Jian, et al.
Veröffentlicht: (2024)
von: Chen, Jian, et al.
Veröffentlicht: (2024)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
von: Yuan, Shenghai, et al.
Veröffentlicht: (2025)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2025)
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
von: Chen, Jian, et al.
Veröffentlicht: (2024)
von: Chen, Jian, et al.
Veröffentlicht: (2024)
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
Self-Evolving 3D Scene Generation from a Single Image
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
HumanNet: Scaling Human-centric Video Learning to One Million Hours
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
Numerical Pruning for Efficient Autoregressive Models
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
Style Customization of Text-to-Vector Generation with Image Diffusion Priors
von: Zhang, Peiying, et al.
Veröffentlicht: (2025)
von: Zhang, Peiying, et al.
Veröffentlicht: (2025)
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
von: Chen, Hong, et al.
Veröffentlicht: (2023)
von: Chen, Hong, et al.
Veröffentlicht: (2023)
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
von: Zhang, Zhehao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2024)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
Output Feedback Control for T‐S Fuzzy Markov Jump Systems Subjected to Parameter‐Dependent Dissipative Performance
von: Jian Wang, et al.
Veröffentlicht: (2024)
von: Jian Wang, et al.
Veröffentlicht: (2024)
Improve Temporal Awareness of LLMs for Sequential Recommendation
von: Chu, Zhendong, et al.
Veröffentlicht: (2024)
von: Chu, Zhendong, et al.
Veröffentlicht: (2024)
AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark
von: Lin, Li, et al.
Veröffentlicht: (2024)
von: Lin, Li, et al.
Veröffentlicht: (2024)
Text-to-Vector Generation with Neural Path Representation
von: Zhang, Peiying, et al.
Veröffentlicht: (2024)
von: Zhang, Peiying, et al.
Veröffentlicht: (2024)
PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
von: Wang, Shulei, et al.
Veröffentlicht: (2025)
von: Wang, Shulei, et al.
Veröffentlicht: (2025)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
von: Flora, James, et al.
Veröffentlicht: (2026)
von: Flora, James, et al.
Veröffentlicht: (2026)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning
von: Zhu, Yihan, et al.
Veröffentlicht: (2025)
von: Zhu, Yihan, et al.
Veröffentlicht: (2025)
Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion
von: Gu, Hengrui, et al.
Veröffentlicht: (2024)
von: Gu, Hengrui, et al.
Veröffentlicht: (2024)
Linear Image Generation by Synthesizing Exposure Brackets
von: Dai, Yuekun, et al.
Veröffentlicht: (2026)
von: Dai, Yuekun, et al.
Veröffentlicht: (2026)
EIT-1M: One Million EEG-Image-Text Pairs for Human Visual-textual Recognition and More
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
von: Liu, Yufan, et al.
Veröffentlicht: (2025)
von: Liu, Yufan, et al.
Veröffentlicht: (2025)
FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation
von: Yao, Zebin, et al.
Veröffentlicht: (2025)
von: Yao, Zebin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner
von: Zhou, Yufan, et al.
Veröffentlicht: (2024) -
Customization Assistant for Text-to-image Generation
von: Zhou, Yufan, et al.
Veröffentlicht: (2023) -
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024) -
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023) -
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)