Benchmarking and Understanding Compositional Relational Reasoning of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Ni, Ruikang, Xiao, Da, Meng, Qingye, Li, Xiangyu, Zheng, Shihui, Liang, Hongliang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improving Transformers with Dynamically Composable Multi-Head Attention
di: Xiao, Da, et al.
Pubblicazione: (2024)
di: Xiao, Da, et al.
Pubblicazione: (2024)
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
di: Xiao, Da, et al.
Pubblicazione: (2025)
di: Xiao, Da, et al.
Pubblicazione: (2025)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
di: Dang, Vy Tuong, et al.
Pubblicazione: (2025)
di: Dang, Vy Tuong, et al.
Pubblicazione: (2025)
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs
di: Liu, Hongliang, et al.
Pubblicazione: (2026)
di: Liu, Hongliang, et al.
Pubblicazione: (2026)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
di: Zhang, Zhihan, et al.
Pubblicazione: (2025)
di: Zhang, Zhihan, et al.
Pubblicazione: (2025)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
di: Dumitran, Adrian-Marius, et al.
Pubblicazione: (2025)
di: Dumitran, Adrian-Marius, et al.
Pubblicazione: (2025)
Probabilistic Reasoning with LLMs for k-anonymity Estimation
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
di: Deng, Jingcheng, et al.
Pubblicazione: (2026)
di: Deng, Jingcheng, et al.
Pubblicazione: (2026)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
di: Yang, Siwei, et al.
Pubblicazione: (2024)
di: Yang, Siwei, et al.
Pubblicazione: (2024)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Benchmarking LLMs' Judgments with No Gold Standard
di: Xu, Shengwei, et al.
Pubblicazione: (2024)
di: Xu, Shengwei, et al.
Pubblicazione: (2024)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
di: Ramezanali, Mohammad, et al.
Pubblicazione: (2025)
di: Ramezanali, Mohammad, et al.
Pubblicazione: (2025)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
di: Wang, Song, et al.
Pubblicazione: (2024)
di: Wang, Song, et al.
Pubblicazione: (2024)
A Decomposition Perspective to Long-context Reasoning for LLMs
di: Xiao, Yanling, et al.
Pubblicazione: (2026)
di: Xiao, Yanling, et al.
Pubblicazione: (2026)
Reasoning Boosts Opinion Alignment in LLMs
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
Learning to Reason in LLMs by Expectation Maximization
di: Lee, Junghyun, et al.
Pubblicazione: (2025)
di: Lee, Junghyun, et al.
Pubblicazione: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
di: Xiao, Feng, et al.
Pubblicazione: (2025)
di: Xiao, Feng, et al.
Pubblicazione: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
di: Li, Xin, et al.
Pubblicazione: (2025)
di: Li, Xin, et al.
Pubblicazione: (2025)
Understanding Knowledge Drift in LLMs through Misinformation
di: Fastowski, Alina, et al.
Pubblicazione: (2024)
di: Fastowski, Alina, et al.
Pubblicazione: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
GeoReasoner: Reasoning On Geospatially Grounded Context For Natural Language Understanding
di: Yan, Yibo, et al.
Pubblicazione: (2024)
di: Yan, Yibo, et al.
Pubblicazione: (2024)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMs
di: Li, Xueyan, et al.
Pubblicazione: (2025)
di: Li, Xueyan, et al.
Pubblicazione: (2025)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
di: Zhao, Guojiang, et al.
Pubblicazione: (2025)
di: Zhao, Guojiang, et al.
Pubblicazione: (2025)
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
Understanding Hidden Computations in Chain-of-Thought Reasoning
di: Bharadwaj, Aryasomayajula Ram
Pubblicazione: (2024)
di: Bharadwaj, Aryasomayajula Ram
Pubblicazione: (2024)
Failure Modes of LLMs for Causal Reasoning on Narratives
di: Yamin, Khurram, et al.
Pubblicazione: (2024)
di: Yamin, Khurram, et al.
Pubblicazione: (2024)
MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
di: Xing, Junjie, et al.
Pubblicazione: (2025)
di: Xing, Junjie, et al.
Pubblicazione: (2025)
Demystifying Long Chain-of-Thought Reasoning in LLMs
di: Yeo, Edward, et al.
Pubblicazione: (2025)
di: Yeo, Edward, et al.
Pubblicazione: (2025)
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding
di: Cheshmi, Seyyed Saeid, et al.
Pubblicazione: (2026)
di: Cheshmi, Seyyed Saeid, et al.
Pubblicazione: (2026)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
di: Hemadri, Raghu Vamshi, et al.
Pubblicazione: (2025)
di: Hemadri, Raghu Vamshi, et al.
Pubblicazione: (2025)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
di: Li, Sihang, et al.
Pubblicazione: (2024)
di: Li, Sihang, et al.
Pubblicazione: (2024)
Few-shot Knowledge Graph Relational Reasoning via Subgraph Adaptation
di: Liu, Haochen, et al.
Pubblicazione: (2024)
di: Liu, Haochen, et al.
Pubblicazione: (2024)
Understanding Memorisation in LLMs: Dynamics, Influencing Factors, and Implications
di: Speicher, Till, et al.
Pubblicazione: (2024)
di: Speicher, Till, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Improving Transformers with Dynamically Composable Multi-Head Attention
di: Xiao, Da, et al.
Pubblicazione: (2024) -
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
di: Xiao, Da, et al.
Pubblicazione: (2025) -
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
di: Lin, Zicheng, et al.
Pubblicazione: (2024) -
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
di: Dang, Vy Tuong, et al.
Pubblicazione: (2025) -
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs
di: Liu, Hongliang, et al.
Pubblicazione: (2026)