MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
Fuente:
arXiv
Salvato in:
| Autori principali: | Wei, Xinyu, Cen, Kangrui, Wei, Hongyang, Guo, Zhen, Cui, Kai, Li, Bairui, Wang, Zeqing, Zhang, Jinrui, Zhang, Lei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
di: Wei, Xinyu, et al.
Pubblicazione: (2025)
di: Wei, Xinyu, et al.
Pubblicazione: (2025)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
di: Wei, Hongyang, et al.
Pubblicazione: (2025)
di: Wei, Hongyang, et al.
Pubblicazione: (2025)
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
di: Qu, Jinyuan, et al.
Pubblicazione: (2026)
di: Qu, Jinyuan, et al.
Pubblicazione: (2026)
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
di: Chen, Zhekai, et al.
Pubblicazione: (2024)
di: Chen, Zhekai, et al.
Pubblicazione: (2024)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
di: Tang, Yin, et al.
Pubblicazione: (2024)
di: Tang, Yin, et al.
Pubblicazione: (2024)
MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
di: Xu, Meng, et al.
Pubblicazione: (2024)
di: Xu, Meng, et al.
Pubblicazione: (2024)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
di: Chen, Zhihong, et al.
Pubblicazione: (2025)
di: Chen, Zhihong, et al.
Pubblicazione: (2025)
CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection
di: Li, Yiheng, et al.
Pubblicazione: (2026)
di: Li, Yiheng, et al.
Pubblicazione: (2026)
Prompt-Free Conditional Diffusion for Multi-object Image Augmentation
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
LayerT2V: A Unified Multi-Layer Video Generation Framework
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
HICH Image/Text (HICH-IT): Comprehensive Text and Image Datasets for Hypertensive Intracerebral Hemorrhage Research
di: Li, Jie, et al.
Pubblicazione: (2024)
di: Li, Jie, et al.
Pubblicazione: (2024)
Compositional Inversion for Stable Diffusion Models
di: Zhang, Xulu, et al.
Pubblicazione: (2023)
di: Zhang, Xulu, et al.
Pubblicazione: (2023)
EarthGPT: A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain
di: Zhang, Wei, et al.
Pubblicazione: (2024)
di: Zhang, Wei, et al.
Pubblicazione: (2024)
Human Image Generation: A Comprehensive Survey
di: Jia, Zhen, et al.
Pubblicazione: (2022)
di: Jia, Zhen, et al.
Pubblicazione: (2022)
EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment
di: Gao, Lancheng, et al.
Pubblicazione: (2026)
di: Gao, Lancheng, et al.
Pubblicazione: (2026)
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
di: Guo, Yuanhe, et al.
Pubblicazione: (2025)
di: Guo, Yuanhe, et al.
Pubblicazione: (2025)
OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion
di: Zhan, Xinyu, et al.
Pubblicazione: (2024)
di: Zhan, Xinyu, et al.
Pubblicazione: (2024)
Advancing Aesthetic Image Generation via Composition Transfer
di: Zou, Kai, et al.
Pubblicazione: (2026)
di: Zou, Kai, et al.
Pubblicazione: (2026)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
CAD 100K: A Comprehensive Multi-Task Dataset for Car Related Visual Anomaly Detection
di: Pang, Jiahua, et al.
Pubblicazione: (2026)
di: Pang, Jiahua, et al.
Pubblicazione: (2026)
Making Images Real Again: A Comprehensive Survey on Deep Image Composition
di: Niu, Li, et al.
Pubblicazione: (2021)
di: Niu, Li, et al.
Pubblicazione: (2021)
UniRef-Image-Edit: Towards Scalable and Consistent Multi-Reference Image Editing
di: Wei, Hongyang, et al.
Pubblicazione: (2026)
di: Wei, Hongyang, et al.
Pubblicazione: (2026)
SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality
di: Lei, Chenyang, et al.
Pubblicazione: (2024)
di: Lei, Chenyang, et al.
Pubblicazione: (2024)
TBI Image/Text (TBI-IT): Comprehensive Text and Image Datasets for Traumatic Brain Injury Research
di: Li, Jie, et al.
Pubblicazione: (2024)
di: Li, Jie, et al.
Pubblicazione: (2024)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
di: Lin, Weifeng, et al.
Pubblicazione: (2025)
di: Lin, Weifeng, et al.
Pubblicazione: (2025)
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
di: Wang, Guanqun, et al.
Pubblicazione: (2024)
di: Wang, Guanqun, et al.
Pubblicazione: (2024)
UNICE: Training A Universal Image Contrast Enhancer
di: Cui, Ruodai, et al.
Pubblicazione: (2025)
di: Cui, Ruodai, et al.
Pubblicazione: (2025)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
di: Cui, Siying, et al.
Pubblicazione: (2024)
di: Cui, Siying, et al.
Pubblicazione: (2024)
SimMAT: Exploring Transferability from Vision Foundation Models to Any Image Modality
di: Lei, Chenyang, et al.
Pubblicazione: (2024)
di: Lei, Chenyang, et al.
Pubblicazione: (2024)
Generative Active Learning for Image Synthesis Personalization
di: Zhang, Xulu, et al.
Pubblicazione: (2024)
di: Zhang, Xulu, et al.
Pubblicazione: (2024)
VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset
di: Chen, Zhizhou, et al.
Pubblicazione: (2026)
di: Chen, Zhizhou, et al.
Pubblicazione: (2026)
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
di: Dong, Xiaoyi, et al.
Pubblicazione: (2024)
di: Dong, Xiaoyi, et al.
Pubblicazione: (2024)
Compositional Attribute Imbalance in Vision Datasets
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling
di: Wei, Hongyang, et al.
Pubblicazione: (2026)
di: Wei, Hongyang, et al.
Pubblicazione: (2026)
OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
di: Li, Hongyang, et al.
Pubblicazione: (2025)
di: Li, Hongyang, et al.
Pubblicazione: (2025)
M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment
di: Cui, Chuan, et al.
Pubblicazione: (2025)
di: Cui, Chuan, et al.
Pubblicazione: (2025)
Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
di: Chen, Jierun, et al.
Pubblicazione: (2024)
di: Chen, Jierun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
di: Wang, Zeqing, et al.
Pubblicazione: (2025) -
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
di: Wei, Xinyu, et al.
Pubblicazione: (2025) -
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
di: Wei, Hongyang, et al.
Pubblicazione: (2025) -
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
di: Qu, Jinyuan, et al.
Pubblicazione: (2026) -
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
di: Chen, Zhekai, et al.
Pubblicazione: (2024)