REAL: Realism Evaluation of Text-to-Image Generation Models for Effective Data Augmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ran, Jin, Xiaomeng, ji, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
by: Pan, Mianzhi, et al.
Published: (2023)
by: Pan, Mianzhi, et al.
Published: (2023)
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
by: Wen, Haokun, et al.
Published: (2024)
by: Wen, Haokun, et al.
Published: (2024)
Contrastive Visual Data Augmentation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
Steganographic Embeddings as an Effective Data Augmentation
by: DiSalvo, Nicholas
Published: (2025)
by: DiSalvo, Nicholas
Published: (2025)
Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data
by: Kumar, Puneet, et al.
Published: (2024)
by: Kumar, Puneet, et al.
Published: (2024)
RFNNS: Robust Fixed Neural Network Steganography with Universal Text-to-Image Models
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
by: Ma, Zhiyong, et al.
Published: (2025)
by: Ma, Zhiyong, et al.
Published: (2025)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
Generating Digital Models Using Text-to-3D and Image-to-3D Prompts: Critical Case Study
by: Ziatdinov, Rushan, et al.
Published: (2025)
by: Ziatdinov, Rushan, et al.
Published: (2025)
Relationship Analysis of Image-Text Pair in SNS Posts
by: Nabeoka, Takuto, et al.
Published: (2025)
by: Nabeoka, Takuto, et al.
Published: (2025)
SPP-SCL: Semi-Push-Pull Supervised Contrastive Learning for Image-Text Sentiment Analysis and Beyond
by: Wu, Jiesheng, et al.
Published: (2026)
by: Wu, Jiesheng, et al.
Published: (2026)
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
by: Xu, Jingning, et al.
Published: (2026)
by: Xu, Jingning, et al.
Published: (2026)
Retrieval-Augmented Text-to-Audio Generation
by: Yuan, Yi, et al.
Published: (2023)
by: Yuan, Yi, et al.
Published: (2023)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Evaluation of Objective Image Quality Metrics for High-Fidelity Image Compression
by: Mohammadi, Shima, et al.
Published: (2025)
by: Mohammadi, Shima, et al.
Published: (2025)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
by: Wang, Hanyao, et al.
Published: (2024)
by: Wang, Hanyao, et al.
Published: (2024)
G-Refine: A General Quality Refiner for Text-to-Image Generation
by: Li, Chunyi, et al.
Published: (2024)
by: Li, Chunyi, et al.
Published: (2024)
On Copyright Risks of Text-to-Image Diffusion Models
by: Zhang, Yang, et al.
Published: (2023)
by: Zhang, Yang, et al.
Published: (2023)
Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image Generation
by: Matsuhira, Chihaya, et al.
Published: (2024)
by: Matsuhira, Chihaya, et al.
Published: (2024)
M3ST-DTI: A multi-task learning model for drug-target interactions based on multi-modal features and multi-stage alignment
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
by: Li, Yaoru, et al.
Published: (2025)
by: Li, Yaoru, et al.
Published: (2025)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated Text
by: Munyer, Travis, et al.
Published: (2023)
by: Munyer, Travis, et al.
Published: (2023)
Development of a Cost‐Effective Data Collection System for Monitoring Urban Traffic Using Automated Intelligent Image Processing and Machine Vision
by: Enkusilasie Tegegne Alemu, et al.
Published: (2026)
by: Enkusilasie Tegegne Alemu, et al.
Published: (2026)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
by: Shi, Junchang, et al.
Published: (2025)
by: Shi, Junchang, et al.
Published: (2025)
An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments
by: Korkiakoski, Mikko, et al.
Published: (2025)
by: Korkiakoski, Mikko, et al.
Published: (2025)
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
by: Liu, Dancheng, et al.
Published: (2025)
by: Liu, Dancheng, et al.
Published: (2025)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
by: Yang, Quanwei, et al.
Published: (2025)
by: Yang, Quanwei, et al.
Published: (2025)
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
by: Gong, Ziyu, et al.
Published: (2024)
by: Gong, Ziyu, et al.
Published: (2024)
Towards Better Text-to-Image Generation Alignment via Attention Modulation
by: Wu, Yihang, et al.
Published: (2024)
by: Wu, Yihang, et al.
Published: (2024)
MusicWeaver: Composer-Style Structural Editing and Minute-Scale Coherent Music Generation
by: Wang, Xuanchen, et al.
Published: (2025)
by: Wang, Xuanchen, et al.
Published: (2025)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
Retrieval-Augmented Multimodal Model for Fake News Detection
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training
by: Ding, Muhe, et al.
Published: (2024)
by: Ding, Muhe, et al.
Published: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
by: Wang, Siyu, et al.
Published: (2025)
by: Wang, Siyu, et al.
Published: (2025)
Retrieval-Augmented Generation for Electrocardiogram-Language Models
by: Song, Xiaoyu, et al.
Published: (2025)
by: Song, Xiaoyu, et al.
Published: (2025)
Similar Items
-
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
by: Pan, Mianzhi, et al.
Published: (2023) -
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
by: Wen, Haokun, et al.
Published: (2024) -
Contrastive Visual Data Augmentation
by: Zhou, Yu, et al.
Published: (2025) -
Steganographic Embeddings as an Effective Data Augmentation
by: DiSalvo, Nicholas
Published: (2025) -
Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data
by: Kumar, Puneet, et al.
Published: (2024)