15M Multimodal Facial Image-Text Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Dawei, Li, YuTang, Liu, YingGe, Jia, Mingming, YuanHui, Zhang, Wang, Guoyin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025)
by: Dai, Dawei, et al.
Published: (2025)
Multi-Granularity Representation Learning for Sketch-based Dynamic Face Image Retrieval
by: Wang, Liang, et al.
Published: (2023)
by: Wang, Liang, et al.
Published: (2023)
Granular-ball Representation Learning for Deep CNN on Learning with Label Noise
by: Dai, Dawei, et al.
Published: (2024)
by: Dai, Dawei, et al.
Published: (2024)
Pore-scale Image Patch Dataset and A Comparative Evaluation of Pore-scale Facial Features
by: Li, Dong, et al.
Published: (2025)
by: Li, Dong, et al.
Published: (2025)
Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025)
by: Dai, Dawei, et al.
Published: (2025)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
by: Zhang, Hui, et al.
Published: (2026)
by: Zhang, Hui, et al.
Published: (2026)
Granular Ball Guided Masking: Structure-aware Data Augmentation
by: Xia, Shuyin, et al.
Published: (2025)
by: Xia, Shuyin, et al.
Published: (2025)
Smile on the Face, Sadness in the Eyes: Bridging the Emotion Gap with a Multimodal Dataset of Eye and Facial Behaviors
by: Liu, Kejun, et al.
Published: (2025)
by: Liu, Kejun, et al.
Published: (2025)
Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
by: Zhang, Fan, et al.
Published: (2025)
by: Zhang, Fan, et al.
Published: (2025)
Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis
by: Zheng, Kaiwen, et al.
Published: (2025)
by: Zheng, Kaiwen, et al.
Published: (2025)
Safety of Multimodal Large Language Models on Images and Texts
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
by: Hu, Zhuozhao, et al.
Published: (2025)
by: Hu, Zhuozhao, et al.
Published: (2025)
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
Spoofing-aware Prompt Learning for Unified Physical-Digital Facial Attack Detection
by: Guo, Jiabao, et al.
Published: (2025)
by: Guo, Jiabao, et al.
Published: (2025)
Anti-Aesthetics: Protecting Facial Privacy against Customized Text-to-Image Synthesis
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
by: Li, Bohao, et al.
Published: (2024)
by: Li, Bohao, et al.
Published: (2024)
TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
by: Wang, Alex Jinpeng, et al.
Published: (2025)
by: Wang, Alex Jinpeng, et al.
Published: (2025)
Anonymization Prompt Learning for Facial Privacy-Preserving Text-to-Image Generation
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
Bridging the Gap: Aligning Text-to-Image Diffusion Models with Specific Feedback
by: Niu, Xuexiang, et al.
Published: (2024)
by: Niu, Xuexiang, et al.
Published: (2024)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
FITA: Fine-grained Image-Text Aligner for Radiology Report Generation
by: Yang, Honglong, et al.
Published: (2024)
by: Yang, Honglong, et al.
Published: (2024)
Brain-Inspired Multimodal Spiking Neural Network for Image-Text Retrieval
by: Zong, Xintao, et al.
Published: (2026)
by: Zong, Xintao, et al.
Published: (2026)
Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models
by: Wang, Luozhou, et al.
Published: (2023)
by: Wang, Luozhou, et al.
Published: (2023)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
by: Yue, Xinli, et al.
Published: (2025)
by: Yue, Xinli, et al.
Published: (2025)
VIP: Versatile Image Outpainting Empowered by Multimodal Large Language Model
by: Yang, Jinze, et al.
Published: (2024)
by: Yang, Jinze, et al.
Published: (2024)
InstaFace: Identity-Preserving Facial Editing with Single Image Inference
by: Khan, MD Wahiduzzaman, et al.
Published: (2025)
by: Khan, MD Wahiduzzaman, et al.
Published: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
Enhancing Single-Image Facial Demorphing using Multimodal Large Language Models
by: Shukla, Nitish, et al.
Published: (2026)
by: Shukla, Nitish, et al.
Published: (2026)
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
by: Lian, Jingchun, et al.
Published: (2024)
by: Lian, Jingchun, et al.
Published: (2024)
SRFlow: A Dataset and Regularization Model for High-Resolution Facial Optical Flow via Splatting Rasterization
by: Zhang, JiaLin, et al.
Published: (2026)
by: Zhang, JiaLin, et al.
Published: (2026)
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification
by: Wang, Jiao, et al.
Published: (2026)
by: Wang, Jiao, et al.
Published: (2026)
Exploring Open-Vocabulary Object Recognition in Images using CLIP
by: Chen, Wei Yu, et al.
Published: (2026)
by: Chen, Wei Yu, et al.
Published: (2026)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models
by: Shuai, Xincheng, et al.
Published: (2024)
by: Shuai, Xincheng, et al.
Published: (2024)
Benchmarking Robustness of Multimodal Image-Text Models under Distribution Shift
by: Qiu, Jielin, et al.
Published: (2022)
by: Qiu, Jielin, et al.
Published: (2022)
Square Superpixel Generation and Representation Learning via Granular Ball Computing
by: Xia, Shuyin, et al.
Published: (2026)
by: Xia, Shuyin, et al.
Published: (2026)
Similar Items
-
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025) -
Multi-Granularity Representation Learning for Sketch-based Dynamic Face Image Retrieval
by: Wang, Liang, et al.
Published: (2023) -
Granular-ball Representation Learning for Deep CNN on Learning with Label Noise
by: Dai, Dawei, et al.
Published: (2024) -
Pore-scale Image Patch Dataset and A Comparative Evaluation of Pore-scale Facial Features
by: Li, Dong, et al.
Published: (2025) -
Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025)