ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Ruiran, Si, Jiasheng, Guo, Zhijiang, Pierrehumbert, Janet B. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decoding Climate Disagreement: A Graph Neural Network-Based Approach to Understanding Social Media Dynamics
von: Su, Ruiran, et al.
Veröffentlicht: (2024)
von: Su, Ruiran, et al.
Veröffentlicht: (2024)
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
von: Shen, Lingdong, et al.
Veröffentlicht: (2024)
von: Shen, Lingdong, et al.
Veröffentlicht: (2024)
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
von: Wang, Siting, et al.
Veröffentlicht: (2025)
von: Wang, Siting, et al.
Veröffentlicht: (2025)
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
ChartBench: A Benchmark for Complex Visual Reasoning in Charts
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2023)
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2023)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
Text Role Classification in Scientific Charts Using Multimodal Transformers
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
von: Tang, Liyan, et al.
Veröffentlicht: (2025)
von: Tang, Liyan, et al.
Veröffentlicht: (2025)
MSG-Chart: Multimodal Scene Graph for ChartQA
von: Dai, Yue, et al.
Veröffentlicht: (2024)
von: Dai, Yue, et al.
Veröffentlicht: (2024)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models
von: Holtermann, Carolin, et al.
Veröffentlicht: (2026)
von: Holtermann, Carolin, et al.
Veröffentlicht: (2026)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
CycleChart: A Unified Consistency-Based Learning Framework for Bidirectional Chart Understanding and Generation
von: Deng, Dazhen, et al.
Veröffentlicht: (2025)
von: Deng, Dazhen, et al.
Veröffentlicht: (2025)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
von: Li, Zhuowan, et al.
Veröffentlicht: (2024)
von: Li, Zhuowan, et al.
Veröffentlicht: (2024)
Multimodal Fact-Level Attribution for Verifiable Reasoning
von: Wan, David, et al.
Veröffentlicht: (2026)
von: Wan, David, et al.
Veröffentlicht: (2026)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations
von: Ford, James, et al.
Veröffentlicht: (2024)
von: Ford, James, et al.
Veröffentlicht: (2024)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2024)
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2024)
On the Perception Bottleneck of VLMs for Chart Understanding
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
CHAOS: Chart Analysis with Outlier Samples
von: Moured, Omar, et al.
Veröffentlicht: (2025)
von: Moured, Omar, et al.
Veröffentlicht: (2025)
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding
von: Zaranis, Emmanouil, et al.
Veröffentlicht: (2025)
von: Zaranis, Emmanouil, et al.
Veröffentlicht: (2025)
Charts Are Not Images: On the Challenges of Scientific Chart Editing
von: Li, Shawn, et al.
Veröffentlicht: (2025)
von: Li, Shawn, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
von: Yang, Yuwei, et al.
Veröffentlicht: (2025)
von: Yang, Yuwei, et al.
Veröffentlicht: (2025)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Decoding Climate Disagreement: A Graph Neural Network-Based Approach to Understanding Social Media Dynamics
von: Su, Ruiran, et al.
Veröffentlicht: (2024) -
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023) -
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
von: Shen, Lingdong, et al.
Veröffentlicht: (2024) -
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024) -
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
von: Wang, Siting, et al.
Veröffentlicht: (2025)