The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wan, Yixin, Wu, Di, Wang, Haoran, Chang, Kai-Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
von: Jha, Akshita, et al.
Veröffentlicht: (2024)
von: Jha, Akshita, et al.
Veröffentlicht: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)
von: Wu, Yin, et al.
Veröffentlicht: (2025)
RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
LLM-RG4: Flexible and Factual Radiology Report Generation across Diverse Input Contexts
von: Wang, Zhuhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhuhao, et al.
Veröffentlicht: (2024)
D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
von: Hickmon, Javon
Veröffentlicht: (2025)
von: Hickmon, Javon
Veröffentlicht: (2025)
Investigating Disability Representations in Text-to-Image Models
von: Tian, Yang, et al.
Veröffentlicht: (2026)
von: Tian, Yang, et al.
Veröffentlicht: (2026)
Auditing Gender Presentation Differences in Text-to-Image Models
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
UniAIDet: A Unified and Universal Benchmark for AI-Generated Image Content Detection and Localization
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency
von: Morbiato, Filippo, et al.
Veröffentlicht: (2025)
von: Morbiato, Filippo, et al.
Veröffentlicht: (2025)
Optimizing Prompts for Text-to-Image Generation
von: Hao, Yaru, et al.
Veröffentlicht: (2022)
von: Hao, Yaru, et al.
Veröffentlicht: (2022)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
LLM-Bootstrapped Targeted Finding Guidance for Factual MLLM-based Medical Report Generation
von: Yang, Cunyuan, et al.
Veröffentlicht: (2026)
von: Yang, Cunyuan, et al.
Veröffentlicht: (2026)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
Generalized People Diversity: Learning a Human Perception-Aligned Diversity Representation for People Images
von: Srinivasan, Hansa, et al.
Veröffentlicht: (2024)
von: Srinivasan, Hansa, et al.
Veröffentlicht: (2024)
CRAFT: Cultural Russian-Oriented Dataset Adaptation for Focused Text-to-Image Generation
von: Vasilev, Viacheslav, et al.
Veröffentlicht: (2025)
von: Vasilev, Viacheslav, et al.
Veröffentlicht: (2025)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories
von: Peterka, Tomas, et al.
Veröffentlicht: (2025)
von: Peterka, Tomas, et al.
Veröffentlicht: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
von: Ashqar, Huthaifa I., et al.
Veröffentlicht: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
Stable Signer: Hierarchical Sign Language Generative Model
von: Fang, Sen, et al.
Veröffentlicht: (2025)
von: Fang, Sen, et al.
Veröffentlicht: (2025)
A Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
von: Ibrahim, Hazem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Hazem, et al.
Veröffentlicht: (2024)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
von: Fraser, Kathleen C., et al.
Veröffentlicht: (2024)
von: Fraser, Kathleen C., et al.
Veröffentlicht: (2024)
Generated Bias: Auditing Internal Bias Dynamics of Text-To-Image Generative Models
von: Mandal, Abhishek, et al.
Veröffentlicht: (2024)
von: Mandal, Abhishek, et al.
Veröffentlicht: (2024)
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
von: Chinchure, Aditya, et al.
Veröffentlicht: (2023)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2023)
Universal Prompt Optimizer for Safe Text-to-Image Generation
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
von: yunhan, Li, et al.
Veröffentlicht: (2025)
von: yunhan, Li, et al.
Veröffentlicht: (2025)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
von: Wu, Di, et al.
Veröffentlicht: (2025) -
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025) -
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects
von: Wan, Yixin, et al.
Veröffentlicht: (2024) -
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
von: Wan, Yixin, et al.
Veröffentlicht: (2024) -
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
von: Wan, Yixin, et al.
Veröffentlicht: (2025)