Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Zijin, Wu, Hao, Dong, Su, Dong, Junnan, Xiao, Yilin, Zhang, Yujing, Wang, Zhu, Huang, Feiran, Li, Linyi, Yang, Hongxia, Huang, Xiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL
von: Hong, Zijin, et al.
Veröffentlicht: (2024)
von: Hong, Zijin, et al.
Veröffentlicht: (2024)
CLR-Bench: Evaluating Large Language Models in College-level Reasoning
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
Structure Guided Large Language Model for SQL Generation
von: Zhang, Qinggang, et al.
Veröffentlicht: (2024)
von: Zhang, Qinggang, et al.
Veröffentlicht: (2024)
Knowledge-to-SQL: Enhancing SQL Generation with Data Expert LLM
von: Hong, Zijin, et al.
Veröffentlicht: (2024)
von: Hong, Zijin, et al.
Veröffentlicht: (2024)
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
Knapsack Optimization-based Schema Linking for LLM-based Text-to-SQL Generation
von: Yuan, Zheng, et al.
Veröffentlicht: (2025)
von: Yuan, Zheng, et al.
Veröffentlicht: (2025)
Cost-efficient Knowledge-based Question Answering with Large Language Models
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
LAG: Logic-Augmented Generation from a Cartesian Perspective
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models
von: Zhang, Qinggang, et al.
Veröffentlicht: (2025)
von: Zhang, Qinggang, et al.
Veröffentlicht: (2025)
ErrorLLM: Modeling SQL Errors for Text-to-SQL Refinement
von: Hong, Zijin, et al.
Veröffentlicht: (2026)
von: Hong, Zijin, et al.
Veröffentlicht: (2026)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Modality-Aware Integration with Large Language Models for Knowledge-based Visual Question Answering
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
von: Dong, Junnan, et al.
Veröffentlicht: (2024)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
Towards Better Question Generation in QA-based Event Extraction
von: Hong, Zijin, et al.
Veröffentlicht: (2024)
von: Hong, Zijin, et al.
Veröffentlicht: (2024)
Use Graph When It Needs: Efficiently and Adaptively Integrating Retrieval-Augmented Generation with Graphs
von: Dong, Su, et al.
Veröffentlicht: (2026)
von: Dong, Su, et al.
Veröffentlicht: (2026)
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
KnowGPT: Knowledge Graph based Prompting for Large Language Models
von: Zhang, Qinggang, et al.
Veröffentlicht: (2023)
von: Zhang, Qinggang, et al.
Veröffentlicht: (2023)
ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
LinearRAG: Linear Graph Retrieval Augmented Generation on Large-scale Corpora
von: Zhuang, Luyao, et al.
Veröffentlicht: (2025)
von: Zhuang, Luyao, et al.
Veröffentlicht: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation
von: Xiang, Zhishang, et al.
Veröffentlicht: (2025)
von: Xiang, Zhishang, et al.
Veröffentlicht: (2025)
Unconstrained Model Merging for Enhanced LLM Reasoning
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Evaluating LLMs' Mathematical Reasoning in Financial Document Question Answering
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
GenDec: A robust generative Question-decomposition method for Multi-hop reasoning
von: Wu, Jian, et al.
Veröffentlicht: (2024)
von: Wu, Jian, et al.
Veröffentlicht: (2024)
Entity Alignment with Noisy Annotations from Large Language Models
von: Chen, Shengyuan, et al.
Veröffentlicht: (2024)
von: Chen, Shengyuan, et al.
Veröffentlicht: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives
von: Liu, Wenxiao, et al.
Veröffentlicht: (2024)
von: Liu, Wenxiao, et al.
Veröffentlicht: (2024)
ChartReasoner: Code-Driven Modality Bridging for Long-Chain Reasoning in Chart Question Answering
von: Jia, Caijun, et al.
Veröffentlicht: (2025)
von: Jia, Caijun, et al.
Veröffentlicht: (2025)
TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering
von: Zhu, Junnan, et al.
Veröffentlicht: (2025)
von: Zhu, Junnan, et al.
Veröffentlicht: (2025)
LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation
von: Xiao, Yilin, et al.
Veröffentlicht: (2026)
von: Xiao, Yilin, et al.
Veröffentlicht: (2026)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
Each Graph is a New Language: Graph Learning with LLMs
von: Zhou, Huachi, et al.
Veröffentlicht: (2025)
von: Zhou, Huachi, et al.
Veröffentlicht: (2025)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Reasoning as State Transition: A Representational Analysis of Reasoning Evolution in Large Language Models
von: Zhang, Siyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Siyuan, et al.
Veröffentlicht: (2026)
STAIR: Improving Safety Alignment with Introspective Reasoning
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL
von: Hong, Zijin, et al.
Veröffentlicht: (2024) -
CLR-Bench: Evaluating Large Language Models in College-level Reasoning
von: Dong, Junnan, et al.
Veröffentlicht: (2024) -
Structure Guided Large Language Model for SQL Generation
von: Zhang, Qinggang, et al.
Veröffentlicht: (2024) -
Knowledge-to-SQL: Enhancing SQL Generation with Data Expert LLM
von: Hong, Zijin, et al.
Veröffentlicht: (2024) -
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)