FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Thomas, Nishal, Thomas, Noel |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale
par: Thomas, Noel
Publié: (2026)
par: Thomas, Noel
Publié: (2026)
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
par: Noël, Valentin
Publié: (2026)
par: Noël, Valentin
Publié: (2026)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
par: Yu, Zhouliang, et autres
Publié: (2025)
par: Yu, Zhouliang, et autres
Publié: (2025)
Lost in Serialization: Invariance and Generalization of LLM Graph Reasoners
par: Herbst, Daniel, et autres
Publié: (2025)
par: Herbst, Daniel, et autres
Publié: (2025)
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
par: Camposampiero, Giacomo, et autres
Publié: (2025)
par: Camposampiero, Giacomo, et autres
Publié: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
par: Hao, Yuren, et autres
Publié: (2025)
par: Hao, Yuren, et autres
Publié: (2025)
xInv: Explainable Optimization of Inverse Problems
par: Memery, Sean, et autres
Publié: (2025)
par: Memery, Sean, et autres
Publié: (2025)
InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees
par: Huang, Chenyu, et autres
Publié: (2026)
par: Huang, Chenyu, et autres
Publié: (2026)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
par: Yao, Jian, et autres
Publié: (2025)
par: Yao, Jian, et autres
Publié: (2025)
DABstep: Data Agent Benchmark for Multi-step Reasoning
par: Egg, Alex, et autres
Publié: (2025)
par: Egg, Alex, et autres
Publié: (2025)
Mathematics and Coding are Universal AI Benchmarks
par: Chojecki, Przemyslaw
Publié: (2025)
par: Chojecki, Przemyslaw
Publié: (2025)
Continuous Invariance Learning
par: Lin, Yong, et autres
Publié: (2023)
par: Lin, Yong, et autres
Publié: (2023)
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
par: Fan, Jingxuan, et autres
Publié: (2024)
par: Fan, Jingxuan, et autres
Publié: (2024)
Cross-Model Semantics in Representation Learning
par: Nikooroo, Saleh, et autres
Publié: (2025)
par: Nikooroo, Saleh, et autres
Publié: (2025)
Bias-Restrained Prefix Representation Finetuning for Mathematical Reasoning
par: Liang, Sirui, et autres
Publié: (2025)
par: Liang, Sirui, et autres
Publié: (2025)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
par: Li, Zhen, et autres
Publié: (2025)
par: Li, Zhen, et autres
Publié: (2025)
Learning with Exact Invariances in Polynomial Time
par: Soleymani, Ashkan, et autres
Publié: (2025)
par: Soleymani, Ashkan, et autres
Publié: (2025)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
par: Zhang, Yuanhe, et autres
Publié: (2025)
par: Zhang, Yuanhe, et autres
Publié: (2025)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
par: Cinquin, Tristan, et autres
Publié: (2025)
par: Cinquin, Tristan, et autres
Publié: (2025)
MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning
par: Lin, Yunze
Publié: (2025)
par: Lin, Yunze
Publié: (2025)
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
par: Pande, Nilay, et autres
Publié: (2025)
par: Pande, Nilay, et autres
Publié: (2025)
PuzzleJAX: A Benchmark for Reasoning and Learning
par: Earle, Sam, et autres
Publié: (2025)
par: Earle, Sam, et autres
Publié: (2025)
PUZZLES: A Benchmark for Neural Algorithmic Reasoning
par: Estermann, Benjamin, et autres
Publié: (2024)
par: Estermann, Benjamin, et autres
Publié: (2024)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
par: Yoa, Seungdong, et autres
Publié: (2026)
par: Yoa, Seungdong, et autres
Publié: (2026)
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
par: Alshammari, Shaden, et autres
Publié: (2026)
par: Alshammari, Shaden, et autres
Publié: (2026)
InvDesFlow-AL: active learning-based workflow for inverse design of functional materials
par: Han, Xiao-Qi, et autres
Publié: (2025)
par: Han, Xiao-Qi, et autres
Publié: (2025)
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
par: Peng, Ruiying, et autres
Publié: (2026)
par: Peng, Ruiying, et autres
Publié: (2026)
Learning Action-based Representations Using Invariance
par: Rudolph, Max, et autres
Publié: (2024)
par: Rudolph, Max, et autres
Publié: (2024)
Handling Distribution Shifts on Graphs: An Invariance Perspective
par: Wu, Qitian, et autres
Publié: (2022)
par: Wu, Qitian, et autres
Publié: (2022)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
par: Mirzadeh, Iman, et autres
Publié: (2024)
par: Mirzadeh, Iman, et autres
Publié: (2024)
CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
par: Zan, Lei, et autres
Publié: (2025)
par: Zan, Lei, et autres
Publié: (2025)
Advanced Weakly-Supervised Formula Exploration for Neuro-Symbolic Mathematical Reasoning
par: Wu, Yuxuan, et autres
Publié: (2025)
par: Wu, Yuxuan, et autres
Publié: (2025)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
par: Lin, Yen-Ting, et autres
Publié: (2025)
par: Lin, Yen-Ting, et autres
Publié: (2025)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
par: Pawar, Pranav, et autres
Publié: (2025)
par: Pawar, Pranav, et autres
Publié: (2025)
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
par: Roggeveen, James V., et autres
Publié: (2025)
par: Roggeveen, James V., et autres
Publié: (2025)
Process In-Context Learning: Enhancing Mathematical Reasoning via Dynamic Demonstration Insertion
par: Gao, Ang, et autres
Publié: (2026)
par: Gao, Ang, et autres
Publié: (2026)
FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis
par: Ondras, Jan, et autres
Publié: (2025)
par: Ondras, Jan, et autres
Publié: (2025)
CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
par: Gao, Zijun, et autres
Publié: (2025)
par: Gao, Zijun, et autres
Publié: (2025)
Operator-Guided Invariance Learning for Continuous Reinforcement Learning
par: Zhang, Zuyuan, et autres
Publié: (2026)
par: Zhang, Zuyuan, et autres
Publié: (2026)
Representation Invariance and Allocation: When Subgroup Balance Matters
par: Alloula, Anissa, et autres
Publié: (2025)
par: Alloula, Anissa, et autres
Publié: (2025)
Documents similaires
-
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale
par: Thomas, Noel
Publié: (2026) -
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
par: Noël, Valentin
Publié: (2026) -
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
par: Yu, Zhouliang, et autres
Publié: (2025) -
Lost in Serialization: Invariance and Generalization of LLM Graph Reasoners
par: Herbst, Daniel, et autres
Publié: (2025) -
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
par: Camposampiero, Giacomo, et autres
Publié: (2025)