More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists
Fuente:
arXiv
Saved in:
| Main Authors: | Alavi, Khashayar, Yeltay, Zhastay, Flek, Lucie, Karimi, Akbar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ArithmAttack: Evaluating Robustness of LLMs to Noisy Context in Math Problem Solving
by: Abedin, Zain Ul, et al.
Published: (2025)
by: Abedin, Zain Ul, et al.
Published: (2025)
Multi-Hop Reasoning for Question Answering with Hyperbolic Representations
by: Welz, Simon, et al.
Published: (2025)
by: Welz, Simon, et al.
Published: (2025)
Improving Low-Resource Dialect Classification Using Retrieval-based Voice Conversion
by: Fischbach, Lea, et al.
Published: (2025)
by: Fischbach, Lea, et al.
Published: (2025)
Exploring Robustness of LLMs to Paraphrasing Based on Sociodemographic Factors
by: Arora, Pulkit, et al.
Published: (2025)
by: Arora, Pulkit, et al.
Published: (2025)
Exploring Robustness of Multilingual LLMs on Real-World Noisy Data
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
Probing the Robustness of Theory of Mind in Large Language Models
by: Nickel, Christian, et al.
Published: (2024)
by: Nickel, Christian, et al.
Published: (2024)
Encoder Fine-tuning with Stochastic Sampling Outperforms Open-weight GPT in Astronomy Knowledge Extraction
by: Rawat, Shivam, et al.
Published: (2025)
by: Rawat, Shivam, et al.
Published: (2025)
Label-Consistent Data Generation for Aspect-Based Sentiment Analysis Using LLM Agents
by: Monfared, Mohammad H. A., et al.
Published: (2026)
by: Monfared, Mohammad H. A., et al.
Published: (2026)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Adversarial Math Word Problem Generation
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?
by: Kao, Kuei-Chun, et al.
Published: (2024)
by: Kao, Kuei-Chun, et al.
Published: (2024)
Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models
by: Nickel, Christian, et al.
Published: (2026)
by: Nickel, Christian, et al.
Published: (2026)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
by: Tong, Yuxuan, et al.
Published: (2024)
by: Tong, Yuxuan, et al.
Published: (2024)
Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
by: Rawat, Shivam, et al.
Published: (2026)
by: Rawat, Shivam, et al.
Published: (2026)
Pitfalls of Conversational LLMs on News Debiasing
by: Schlicht, Ipek Baris, et al.
Published: (2024)
by: Schlicht, Ipek Baris, et al.
Published: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
by: Fang, Meng, et al.
Published: (2024)
by: Fang, Meng, et al.
Published: (2024)
Improving Coherence and Persistence in Agentic AI for System Optimization
by: Karimi, Pantea, et al.
Published: (2026)
by: Karimi, Pantea, et al.
Published: (2026)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
by: Colle, Vincenzo, et al.
Published: (2025)
by: Colle, Vincenzo, et al.
Published: (2025)
A Diversity-Enhanced Knowledge Distillation Model for Practical Math Word Problem Solving
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model
by: Yang, Zhen, et al.
Published: (2024)
by: Yang, Zhen, et al.
Published: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis
by: Yu, Zixiong, et al.
Published: (2026)
by: Yu, Zixiong, et al.
Published: (2026)
Can LLM Agents Identify Spoken Dialects like a Linguist?
by: Bystrich, Tobias, et al.
Published: (2026)
by: Bystrich, Tobias, et al.
Published: (2026)
REAMS: Reasoning Enhanced Algorithm for Maths Solving
by: Singh, Eishkaran, et al.
Published: (2025)
by: Singh, Eishkaran, et al.
Published: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
by: Kurz, Simon, et al.
Published: (2024)
by: Kurz, Simon, et al.
Published: (2024)
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
by: Fatimah, Shiza, et al.
Published: (2026)
by: Fatimah, Shiza, et al.
Published: (2026)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
Modular Arithmetic: Language Models Solve Math Digit by Digit
by: Baeumel, Tanja, et al.
Published: (2025)
by: Baeumel, Tanja, et al.
Published: (2025)
STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
by: Wu, Jiaao, et al.
Published: (2026)
by: Wu, Jiaao, et al.
Published: (2026)
Can Stories Help LLMs Reason? Curating Information Space Through Narrative
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
Bridging Information Gaps with Comprehensive Answers: Improving the Diversity and Informativeness of Follow-Up Questions
by: Liu, Zhe, et al.
Published: (2025)
by: Liu, Zhe, et al.
Published: (2025)
Tucano 2 Cool: Better Open Source LLMs for Portuguese
by: Corrêa, Nicholas Kluge, et al.
Published: (2026)
by: Corrêa, Nicholas Kluge, et al.
Published: (2026)
Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
by: Renze, Matthew, et al.
Published: (2024)
by: Renze, Matthew, et al.
Published: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
USDC: A Dataset of $\underline{U}$ser $\underline{S}$tance and $\underline{D}$ogmatism in Long $\underline{C}$onversations
by: Marreddy, Mounika, et al.
Published: (2024)
by: Marreddy, Mounika, et al.
Published: (2024)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
by: Era, Jalisha Jashim, et al.
Published: (2025)
by: Era, Jalisha Jashim, et al.
Published: (2025)
Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
by: Tang, Xiangru, et al.
Published: (2025)
by: Tang, Xiangru, et al.
Published: (2025)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
by: Gao, Songyang, et al.
Published: (2025)
by: Gao, Songyang, et al.
Published: (2025)
MapCoder: Multi-Agent Code Generation for Competitive Problem Solving
by: Islam, Md. Ashraful, et al.
Published: (2024)
by: Islam, Md. Ashraful, et al.
Published: (2024)
Similar Items
-
ArithmAttack: Evaluating Robustness of LLMs to Noisy Context in Math Problem Solving
by: Abedin, Zain Ul, et al.
Published: (2025) -
Multi-Hop Reasoning for Question Answering with Hyperbolic Representations
by: Welz, Simon, et al.
Published: (2025) -
Improving Low-Resource Dialect Classification Using Retrieval-based Voice Conversion
by: Fischbach, Lea, et al.
Published: (2025) -
Exploring Robustness of LLMs to Paraphrasing Based on Sociodemographic Factors
by: Arora, Pulkit, et al.
Published: (2025) -
Exploring Robustness of Multilingual LLMs on Real-World Noisy Data
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)