Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Xingyu, He, Muyu, Lu, Yujie, Wang, William Yang, Roth, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2023)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2023)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
von: Chen, Pei-An, et al.
Veröffentlicht: (2026)
von: Chen, Pei-An, et al.
Veröffentlicht: (2026)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
Commonsense for Zero-Shot Natural Language Video Localization
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
von: Zhang, Zhi, et al.
Veröffentlicht: (2023)
von: Zhang, Zhi, et al.
Veröffentlicht: (2023)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
VCD: A Dataset for Visual Commonsense Discovery in Images
von: Shen, Xiangqing, et al.
Veröffentlicht: (2024)
von: Shen, Xiangqing, et al.
Veröffentlicht: (2024)
BLINK: Multimodal Large Language Models Can See but Not Perceive
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
von: Park, Jun-Hyung, et al.
Veröffentlicht: (2024)
von: Park, Jun-Hyung, et al.
Veröffentlicht: (2024)
Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
Augmented Commonsense Knowledge for Remote Object Grounding
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
Teaching Text-to-Image Models to Communicate in Dialog
von: Sun, Xiaowen, et al.
Veröffentlicht: (2023)
von: Sun, Xiaowen, et al.
Veröffentlicht: (2023)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
von: Sanli, Enes, et al.
Veröffentlicht: (2025)
von: Sanli, Enes, et al.
Veröffentlicht: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
von: Ma, Mingjie, et al.
Veröffentlicht: (2024)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
A Study of Commonsense Reasoning over Visual Object Properties
von: Kolari, Abhishek, et al.
Veröffentlicht: (2025)
von: Kolari, Abhishek, et al.
Veröffentlicht: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
Enhancing Scene Graph Generation with Hierarchical Relationships and Commonsense Knowledge
von: Jiang, Bowen, et al.
Veröffentlicht: (2023)
von: Jiang, Bowen, et al.
Veröffentlicht: (2023)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
von: Wang, Eileen, et al.
Veröffentlicht: (2024)
Auditing Gender Presentation Differences in Text-to-Image Models
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
Can Vision Language Models Understand Mimed Actions?
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2023) -
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023) -
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
von: Chen, Jiali, et al.
Veröffentlicht: (2024) -
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
von: Li, Xiujun, et al.
Veröffentlicht: (2023) -
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)