ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Liyan, Kim, Grace, Zhao, Xinyu, Lake, Thom, Ding, Wenxuan, Yin, Fangcong, Singhal, Prasann, Wadhwa, Manya, Liu, Zeyu Leo, Sprague, Zayne, Namuduri, Ramya, Hu, Bodun, Rodriguez, Juan Diego, Peng, Puyuan, Durrett, Greg |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024)
by: Sprague, Zayne, et al.
Published: (2024)
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
by: Sprague, Zayne, et al.
Published: (2025)
by: Sprague, Zayne, et al.
Published: (2025)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
by: Namuduri, Ramya, et al.
Published: (2025)
by: Namuduri, Ramya, et al.
Published: (2025)
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024)
by: Wadhwa, Manya, et al.
Published: (2024)
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
A Long Way to Go: Investigating Length Correlations in RLHF
by: Singhal, Prasann, et al.
Published: (2023)
by: Singhal, Prasann, et al.
Published: (2023)
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
by: Lake, Thom, et al.
Published: (2024)
by: Lake, Thom, et al.
Published: (2024)
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
by: Sprague, Zayne, et al.
Published: (2023)
by: Sprague, Zayne, et al.
Published: (2023)
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
by: Tripathi, Tuhina, et al.
Published: (2025)
by: Tripathi, Tuhina, et al.
Published: (2025)
D2PO: Discriminator-Guided DPO with Response Evaluation Models
by: Singhal, Prasann, et al.
Published: (2024)
by: Singhal, Prasann, et al.
Published: (2024)
Using Natural Language Explanations to Rescale Human Judgments
by: Wadhwa, Manya, et al.
Published: (2023)
by: Wadhwa, Manya, et al.
Published: (2023)
Learning Composable Chains-of-Thought
by: Yin, Fangcong, et al.
Published: (2025)
by: Yin, Fangcong, et al.
Published: (2025)
LoFiT: Localized Fine-tuning on LLM Representations
by: Yin, Fangcong, et al.
Published: (2024)
by: Yin, Fangcong, et al.
Published: (2024)
CREATE: Testing LLMs for Associative Creativity
by: Wadhwa, Manya, et al.
Published: (2026)
by: Wadhwa, Manya, et al.
Published: (2026)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
by: Liu, Zeyu Leo, et al.
Published: (2025)
by: Liu, Zeyu Leo, et al.
Published: (2025)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
by: Ding, Wenxuan, et al.
Published: (2026)
by: Ding, Wenxuan, et al.
Published: (2026)
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
by: Subbiah, Melanie, et al.
Published: (2025)
by: Subbiah, Melanie, et al.
Published: (2025)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
by: Liu, Zeyu Leo, et al.
Published: (2024)
by: Liu, Zeyu Leo, et al.
Published: (2024)
RankAlign: A Ranking View of the Generator-Validator Gap in Large Language Models
by: Rodriguez, Juan Diego, et al.
Published: (2025)
by: Rodriguez, Juan Diego, et al.
Published: (2025)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
by: Gunjal, Anisha, et al.
Published: (2024)
by: Gunjal, Anisha, et al.
Published: (2024)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
by: Ye, Xi, et al.
Published: (2025)
by: Ye, Xi, et al.
Published: (2025)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
by: Wang, Songtao, et al.
Published: (2026)
by: Wang, Songtao, et al.
Published: (2026)
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
by: Jiang, Yuru, et al.
Published: (2025)
by: Jiang, Yuru, et al.
Published: (2025)
Complex Claim Verification with Evidence Retrieved in the Wild
by: Chen, Jifan, et al.
Published: (2023)
by: Chen, Jifan, et al.
Published: (2023)
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing
by: Wei, Jiayi, et al.
Published: (2023)
by: Wei, Jiayi, et al.
Published: (2023)
Assessing Robustness to Spurious Correlations in Post-Training Language Models
by: Shuieh, Julia, et al.
Published: (2025)
by: Shuieh, Julia, et al.
Published: (2025)
X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs
by: Rodriguez, Juan Diego, et al.
Published: (2023)
by: Rodriguez, Juan Diego, et al.
Published: (2023)
Next‐Generation Approach and Mechanistic Insight Mediated Beneficial Plant‐Microbe Interactions to Foster Resilient Agroecosystems and Sustain Soil Health
by: Sudhir Kumar Upadhyay, et al.
Published: (2026)
by: Sudhir Kumar Upadhyay, et al.
Published: (2026)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
by: Xin, Yutong, et al.
Published: (2026)
by: Xin, Yutong, et al.
Published: (2026)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
by: Sriram, Aniruddh, et al.
Published: (2024)
by: Sriram, Aniruddh, et al.
Published: (2024)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
Unusual properties of contact processes on percolated graphs
by: Durrett, Rick
Published: (2024)
by: Durrett, Rick
Published: (2024)
Probability : theory and examples / Rick Durrett
by: Durrett, Rick
by: Durrett, Rick
Expanding the Meaning of "Bride": Arunkumar and Sreeja v Inspector General of Registration (2019)
by: Manya Goel
Published: (2026)
by: Manya Goel
Published: (2026)
Técnicas de resolución de problemas de satisfacción de restricciones
by: Felip Manyà
Published: (2003)
by: Felip Manyà
Published: (2003)
Similar Items
-
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024) -
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
by: Sprague, Zayne, et al.
Published: (2025) -
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025) -
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
by: Namuduri, Ramya, et al.
Published: (2025) -
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024)