TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Wenhao, Yang, Yi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
por: Wang, Wenhao, et al.
Publicado: (2025)
por: Wang, Wenhao, et al.
Publicado: (2025)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
por: Wang, Wenhao, et al.
Publicado: (2024)
por: Wang, Wenhao, et al.
Publicado: (2024)
Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
por: Zhou, Yufan, et al.
Publicado: (2024)
por: Zhou, Yufan, et al.
Publicado: (2024)
TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts
por: Zhuang, Jingyu, et al.
Publicado: (2024)
por: Zhuang, Jingyu, et al.
Publicado: (2024)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
por: Ren, Weiming, et al.
Publicado: (2024)
por: Ren, Weiming, et al.
Publicado: (2024)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
por: Wang, Yuhan, et al.
Publicado: (2025)
por: Wang, Yuhan, et al.
Publicado: (2025)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
por: Yuan, Shenghai, et al.
Publicado: (2025)
por: Yuan, Shenghai, et al.
Publicado: (2025)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
por: Fang, Rongyao, et al.
Publicado: (2025)
por: Fang, Rongyao, et al.
Publicado: (2025)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
por: Wang, Wenhao, et al.
Publicado: (2025)
por: Wang, Wenhao, et al.
Publicado: (2025)
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
por: Flora, James, et al.
Publicado: (2026)
por: Flora, James, et al.
Publicado: (2026)
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
por: Li, Teng, et al.
Publicado: (2025)
por: Li, Teng, et al.
Publicado: (2025)
Fast Prompt Alignment for Text-to-Image Generation
por: Mrini, Khalil, et al.
Publicado: (2024)
por: Mrini, Khalil, et al.
Publicado: (2024)
GuardT2I: Defending Text-to-Image Models from Adversarial Prompts
por: Yang, Yijun, et al.
Publicado: (2024)
por: Yang, Yijun, et al.
Publicado: (2024)
Prompt Refinement with Image Pivot for Text-to-Image Generation
por: Zhan, Jingtao, et al.
Publicado: (2024)
por: Zhan, Jingtao, et al.
Publicado: (2024)
TIP: Tabular-Image Pre-training for Multimodal Classification with Incomplete Data
por: Du, Siyi, et al.
Publicado: (2024)
por: Du, Siyi, et al.
Publicado: (2024)
PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation
por: Jing, Zonglei, et al.
Publicado: (2025)
por: Jing, Zonglei, et al.
Publicado: (2025)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
por: Shi, Xiaoyu, et al.
Publicado: (2024)
por: Shi, Xiaoyu, et al.
Publicado: (2024)
TIPO: Text to Image with Text Presampling for Prompt Optimization
por: Yeh, Shih-Ying, et al.
Publicado: (2024)
por: Yeh, Shih-Ying, et al.
Publicado: (2024)
DeMamba: AI-Generated Video Detection on Million-Scale GenVideo Benchmark
por: Chen, Haoxing, et al.
Publicado: (2024)
por: Chen, Haoxing, et al.
Publicado: (2024)
Towards Safe Synthetic Image Generation On the Web: A Multimodal Robust NSFW Defense and Million Scale Dataset
por: Muneer, Muhammad Shahid, et al.
Publicado: (2025)
por: Muneer, Muhammad Shahid, et al.
Publicado: (2025)
Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing
por: Yang, Yitong, et al.
Publicado: (2024)
por: Yang, Yitong, et al.
Publicado: (2024)
Optimizing Prompts for Text-to-Image Generation
por: Hao, Yaru, et al.
Publicado: (2022)
por: Hao, Yaru, et al.
Publicado: (2022)
TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
por: Wu, Xiangyu, et al.
Publicado: (2024)
por: Wu, Xiangyu, et al.
Publicado: (2024)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
por: Guo, Xun, et al.
Publicado: (2023)
por: Guo, Xun, et al.
Publicado: (2023)
Universal Prompt Optimizer for Safe Text-to-Image Generation
por: Wu, Zongyu, et al.
Publicado: (2024)
por: Wu, Zongyu, et al.
Publicado: (2024)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
por: Wang, Cong, et al.
Publicado: (2023)
por: Wang, Cong, et al.
Publicado: (2023)
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
por: Liu, Peng, et al.
Publicado: (2025)
por: Liu, Peng, et al.
Publicado: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
por: Gao, Jiayi, et al.
Publicado: (2025)
por: Gao, Jiayi, et al.
Publicado: (2025)
Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs
por: Chen, Zhi-Kai, et al.
Publicado: (2026)
por: Chen, Zhi-Kai, et al.
Publicado: (2026)
AnimeDL-2M: Million-Scale AI-Generated Anime Image Detection and Localization in Diffusion Era
por: Zhu, Chenyang, et al.
Publicado: (2025)
por: Zhu, Chenyang, et al.
Publicado: (2025)
ProTIP: Probabilistic Robustness Verification on Text-to-Image Diffusion Models against Stochastic Perturbation
por: Zhang, Yi, et al.
Publicado: (2024)
por: Zhang, Yi, et al.
Publicado: (2024)
ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
por: Krumpl, Gerhard, et al.
Publicado: (2026)
por: Krumpl, Gerhard, et al.
Publicado: (2026)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
por: Yang, Xiangpeng, et al.
Publicado: (2024)
por: Yang, Xiangpeng, et al.
Publicado: (2024)
VersaT2I: Improving Text-to-Image Models with Versatile Reward
por: Guo, Jianshu, et al.
Publicado: (2024)
por: Guo, Jianshu, et al.
Publicado: (2024)
Quality and Quantity: Unveiling a Million High-Quality Images for Text-to-Image Synthesis in Fashion Design
por: Yu, Jia, et al.
Publicado: (2023)
por: Yu, Jia, et al.
Publicado: (2023)
AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark
por: Lin, Li, et al.
Publicado: (2024)
por: Lin, Li, et al.
Publicado: (2024)
TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
por: Wang, Alex Jinpeng, et al.
Publicado: (2025)
por: Wang, Alex Jinpeng, et al.
Publicado: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
por: Liu, Yexin, et al.
Publicado: (2025)
por: Liu, Yexin, et al.
Publicado: (2025)
Distill Video Datasets into Images
por: Zhao, Zhenghao, et al.
Publicado: (2025)
por: Zhao, Zhenghao, et al.
Publicado: (2025)
Iterative Prompt Refinement for Safer Text-to-Image Generation
por: Jeon, Jinwoo, et al.
Publicado: (2025)
por: Jeon, Jinwoo, et al.
Publicado: (2025)
Ejemplares similares
-
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
por: Wang, Wenhao, et al.
Publicado: (2025) -
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
por: Wang, Wenhao, et al.
Publicado: (2024) -
Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
por: Zhou, Yufan, et al.
Publicado: (2024) -
TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts
por: Zhuang, Jingyu, et al.
Publicado: (2024) -
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
por: Ren, Weiming, et al.
Publicado: (2024)