ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Leixin, Eger, Steffen, Cheng, Yinjie, Zhai, Weihe, Belouadi, Jonas, Leiter, Christoph, Ponzetto, Simone Paolo, Moafian, Fahimeh, Zhao, Zhixue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
by: Belouadi, Jonas, et al.
Published: (2024)
by: Belouadi, Jonas, et al.
Published: (2024)
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
by: Belouadi, Jonas, et al.
Published: (2023)
by: Belouadi, Jonas, et al.
Published: (2023)
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
by: Belouadi, Jonas, et al.
Published: (2022)
by: Belouadi, Jonas, et al.
Published: (2022)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
by: Belouadi, Jonas, et al.
Published: (2022)
by: Belouadi, Jonas, et al.
Published: (2022)
TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
by: Belouadi, Jonas, et al.
Published: (2025)
by: Belouadi, Jonas, et al.
Published: (2025)
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
BMX: Boosting Natural Language Generation Metrics with Explainability
by: Leiter, Christoph, et al.
Published: (2022)
by: Leiter, Christoph, et al.
Published: (2022)
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
by: Leiter, Christoph, et al.
Published: (2025)
by: Leiter, Christoph, et al.
Published: (2025)
Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation
by: Zhang, Ran, et al.
Published: (2026)
by: Zhang, Ran, et al.
Published: (2026)
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
by: Kiefer, Lotta, et al.
Published: (2026)
by: Kiefer, Lotta, et al.
Published: (2026)
MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models
by: Belouadi, Jonas, et al.
Published: (2025)
by: Belouadi, Jonas, et al.
Published: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
Towards Explainable Evaluation Metrics for Machine Translation
by: Leiter, Christoph, et al.
Published: (2023)
by: Leiter, Christoph, et al.
Published: (2023)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
by: Larionov, Daniil, et al.
Published: (2025)
by: Larionov, Daniil, et al.
Published: (2025)
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
by: Greisinger, Christian, et al.
Published: (2026)
by: Greisinger, Christian, et al.
Published: (2026)
ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications
by: Takeshita, Sotaro, et al.
Published: (2024)
by: Takeshita, Sotaro, et al.
Published: (2024)
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
by: Wang, Zhipin, et al.
Published: (2026)
by: Wang, Zhipin, et al.
Published: (2026)
To MRL or not to MRL: Text Embeddings are Robust to Truncation Without Matryoshka Learning, Except In Heavy Truncation Scenarios
by: Takeshita, Sotaro, et al.
Published: (2026)
by: Takeshita, Sotaro, et al.
Published: (2026)
Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs
by: Bombieri, Marco, et al.
Published: (2026)
by: Bombieri, Marco, et al.
Published: (2026)
ROUGE-K: Do Your Summaries Have Keywords?
by: Takeshita, Sotaro, et al.
Published: (2024)
by: Takeshita, Sotaro, et al.
Published: (2024)
Enriching Social Science Research via Survey Item Linking
by: Tsereteli, Tornike, et al.
Published: (2024)
by: Tsereteli, Tornike, et al.
Published: (2024)
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
by: Ramesh, Samarth N, et al.
Published: (2024)
by: Ramesh, Samarth N, et al.
Published: (2024)
Do LLMs Dream of Ontologies?
by: Bombieri, Marco, et al.
Published: (2024)
by: Bombieri, Marco, et al.
Published: (2024)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
by: Zhai, Weihe, et al.
Published: (2023)
by: Zhai, Weihe, et al.
Published: (2023)
PET: An Annotated Dataset for Process Extraction from Natural Language Text
by: Bellan, Patrizio, et al.
Published: (2022)
by: Bellan, Patrizio, et al.
Published: (2022)
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
by: Roy, Subhadeep, et al.
Published: (2026)
by: Roy, Subhadeep, et al.
Published: (2026)
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks
by: Chen, Yanran, et al.
Published: (2025)
by: Chen, Yanran, et al.
Published: (2025)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
by: Larionov, Daniil, et al.
Published: (2024)
by: Larionov, Daniil, et al.
Published: (2024)
LLM-based multi-agent poetry generation in non-cooperative environments
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
by: Larionov, Daniil, et al.
Published: (2025)
by: Larionov, Daniil, et al.
Published: (2025)
Is there really a Citation Age Bias in NLP?
by: Nguyen, Hoa, et al.
Published: (2024)
by: Nguyen, Hoa, et al.
Published: (2024)
3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
by: Ahmed, Noor, et al.
Published: (2025)
by: Ahmed, Noor, et al.
Published: (2025)
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
by: Zhang, Ran, et al.
Published: (2023)
by: Zhang, Ran, et al.
Published: (2023)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
by: Cheng, Yinjie, et al.
Published: (2025)
by: Cheng, Yinjie, et al.
Published: (2025)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
by: Chen, Yuxin, et al.
Published: (2024)
by: Chen, Yuxin, et al.
Published: (2024)
A Review on Generative AI For Text-To-Image and Image-To-Image Generation and Implications To Scientific Images
by: Sordo, Zineb, et al.
Published: (2025)
by: Sordo, Zineb, et al.
Published: (2025)
Evaluating Diversity in Automatic Poetry Generation
by: Chen, Yanran, et al.
Published: (2024)
by: Chen, Yanran, et al.
Published: (2024)
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
by: Klerings, Alina, et al.
Published: (2025)
by: Klerings, Alina, et al.
Published: (2025)
Similar Items
-
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
by: Belouadi, Jonas, et al.
Published: (2024) -
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
by: Belouadi, Jonas, et al.
Published: (2023) -
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
by: Belouadi, Jonas, et al.
Published: (2022) -
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
by: Belouadi, Jonas, et al.
Published: (2022) -
TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
by: Belouadi, Jonas, et al.
Published: (2025)