Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Junling, Chen, Boqi, Do, Heejin, Akhtar, Mubashara, Wang, April Yi, Sachan, Mrinmaya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models
von: Wang, Junling, et al.
Veröffentlicht: (2025)
von: Wang, Junling, et al.
Veröffentlicht: (2025)
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
Do Vision-Language Models Really Understand Visual Language?
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
von: Do, Heejin, et al.
Veröffentlicht: (2026)
von: Do, Heejin, et al.
Veröffentlicht: (2026)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
von: Pang, Xingzhou, et al.
Veröffentlicht: (2026)
von: Pang, Xingzhou, et al.
Veröffentlicht: (2026)
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
von: Wang, Yucheng, et al.
Veröffentlicht: (2025)
von: Wang, Yucheng, et al.
Veröffentlicht: (2025)
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
von: Chi, Ziheng, et al.
Veröffentlicht: (2025)
von: Chi, Ziheng, et al.
Veröffentlicht: (2025)
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)
von: Wu, Yin, et al.
Veröffentlicht: (2025)
Book2Dial: Generating Teacher-Student Interactions from Textbooks for Cost-Effective Development of Educational Chatbots
von: Wang, Junling, et al.
Veröffentlicht: (2024)
von: Wang, Junling, et al.
Veröffentlicht: (2024)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
Emergent Visual-Semantic Hierarchies in Image-Text Representations
von: Alper, Morris, et al.
Veröffentlicht: (2024)
von: Alper, Morris, et al.
Veröffentlicht: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
Probing for Arithmetic Errors in Language Models
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
von: Gao, Xin, et al.
Veröffentlicht: (2026)
von: Gao, Xin, et al.
Veröffentlicht: (2026)
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2023)
Universal Prompt Optimizer for Safe Text-to-Image Generation
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
von: Wu, Zongyu, et al.
Veröffentlicht: (2024)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
Bilingual Text-to-Motion Generation: A New Benchmark and Baselines
von: Weng, Wanjiang, et al.
Veröffentlicht: (2026)
von: Weng, Wanjiang, et al.
Veröffentlicht: (2026)
Fast Prompt Alignment for Text-to-Image Generation
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
Multi-Head Attention Driven Dynamic Visual-Semantic Embedding for Enhanced Image-Text Matching
von: Chen, Wenjing
Veröffentlicht: (2024)
von: Chen, Wenjing
Veröffentlicht: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
von: Zhao, Yi, et al.
Veröffentlicht: (2026)
von: Zhao, Yi, et al.
Veröffentlicht: (2026)
The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation
von: He, Yuze, et al.
Veröffentlicht: (2023)
von: He, Yuze, et al.
Veröffentlicht: (2023)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
EVA-02: A Visual Representation for Neon Genesis
von: Fang, Yuxin, et al.
Veröffentlicht: (2023)
von: Fang, Yuxin, et al.
Veröffentlicht: (2023)
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
von: Vasilev, Viacheslav, et al.
Veröffentlicht: (2025)
von: Vasilev, Viacheslav, et al.
Veröffentlicht: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
von: Black, Alexander, et al.
Veröffentlicht: (2024)
von: Black, Alexander, et al.
Veröffentlicht: (2024)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
von: Jha, Akshita, et al.
Veröffentlicht: (2024)
von: Jha, Akshita, et al.
Veröffentlicht: (2024)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
von: Wang, Ziteng, et al.
Veröffentlicht: (2026)
von: Wang, Ziteng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models
von: Wang, Junling, et al.
Veröffentlicht: (2025) -
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025) -
Do Vision-Language Models Really Understand Visual Language?
von: Hou, Yifan, et al.
Veröffentlicht: (2024) -
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
von: Do, Heejin, et al.
Veröffentlicht: (2026) -
Unveiling the Visual Counting Bottleneck in Vision-Language Models
von: Pang, Xingzhou, et al.
Veröffentlicht: (2026)