Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Jiayu, Ming, Yifei, Ke, Zixuan, Xiong, Caiming, Joty, Shafiq, Albarghouthi, Aws, Sala, Frederic |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SkillOrchestra: Learning to Route Agents via Skill Transfer
di: Wang, Jiayu, et al.
Pubblicazione: (2026)
di: Wang, Jiayu, et al.
Pubblicazione: (2026)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
Demystifying Domain-adaptive Post-training for Financial LLMs
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
NAACL2025 Tutorial: Adaptation of Large Language Models
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
di: Ming, Yifei, et al.
Pubblicazione: (2025)
di: Ming, Yifei, et al.
Pubblicazione: (2025)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
di: Ke, Zixuan, et al.
Pubblicazione: (2026)
di: Ke, Zixuan, et al.
Pubblicazione: (2026)
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
di: Ming, Yifei, et al.
Pubblicazione: (2024)
di: Ming, Yifei, et al.
Pubblicazione: (2024)
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2024)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2024)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2025)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
di: Niu, Tong, et al.
Pubblicazione: (2024)
di: Niu, Tong, et al.
Pubblicazione: (2024)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Harnessing LLM Agents with Skill Programs
di: Liu, Hongjun, et al.
Pubblicazione: (2026)
di: Liu, Hongjun, et al.
Pubblicazione: (2026)
Modeling Uncertainty and Using Post-fusion as Fallback Improves Retrieval Augmented Generation with LLMs
di: Liu, Ye, et al.
Pubblicazione: (2023)
di: Liu, Ye, et al.
Pubblicazione: (2023)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2026)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2026)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
di: Shi, Zhenmei, et al.
Pubblicazione: (2024)
di: Shi, Zhenmei, et al.
Pubblicazione: (2024)
Efficiently Aligned Cross-Lingual Transfer Learning for Conversational Tasks using Prompt-Tuning
di: Tu, Lifu, et al.
Pubblicazione: (2023)
di: Tu, Lifu, et al.
Pubblicazione: (2023)
Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
di: Tarunokusumo, Ravindra Aribowo, et al.
Pubblicazione: (2025)
di: Tarunokusumo, Ravindra Aribowo, et al.
Pubblicazione: (2025)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
di: Liu, Ye, et al.
Pubblicazione: (2024)
di: Liu, Ye, et al.
Pubblicazione: (2024)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Linear-Time T-Gate Optimization via Random Abstraction
di: Albarghouthi, Aws
Pubblicazione: (2026)
di: Albarghouthi, Aws
Pubblicazione: (2026)
Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
di: Jiao, Fangkai, et al.
Pubblicazione: (2024)
di: Jiao, Fangkai, et al.
Pubblicazione: (2024)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
di: Orlanski, Gabriel, et al.
Pubblicazione: (2026)
di: Orlanski, Gabriel, et al.
Pubblicazione: (2026)
Pareto Optimal Code Generation
di: Orlanski, Gabriel, et al.
Pubblicazione: (2025)
di: Orlanski, Gabriel, et al.
Pubblicazione: (2025)
U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning
di: Lee, Christine P, et al.
Pubblicazione: (2026)
di: Lee, Christine P, et al.
Pubblicazione: (2026)
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning
di: Yang, Xia, et al.
Pubblicazione: (2026)
di: Yang, Xia, et al.
Pubblicazione: (2026)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
di: Guo, Xiaojun, et al.
Pubblicazione: (2025)
di: Guo, Xiaojun, et al.
Pubblicazione: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
di: Li, Yuanhao, et al.
Pubblicazione: (2025)
di: Li, Yuanhao, et al.
Pubblicazione: (2025)
Dissecting the Failure of Invariant Learning on Graphs
di: Wang, Qixun, et al.
Pubblicazione: (2024)
di: Wang, Qixun, et al.
Pubblicazione: (2024)
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
di: Yoshihara, Hiroshi, et al.
Pubblicazione: (2025)
di: Yoshihara, Hiroshi, et al.
Pubblicazione: (2025)
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
di: Han, Simeng, et al.
Pubblicazione: (2024)
di: Han, Simeng, et al.
Pubblicazione: (2024)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
di: Venkataramani, Vishal, et al.
Pubblicazione: (2026)
di: Venkataramani, Vishal, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SkillOrchestra: Learning to Route Agents via Skill Transfer
di: Wang, Jiayu, et al.
Pubblicazione: (2026) -
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
di: Wang, Jiayu, et al.
Pubblicazione: (2025) -
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
di: Wang, Jiayu, et al.
Pubblicazione: (2025) -
Demystifying Domain-adaptive Post-training for Financial LLMs
di: Ke, Zixuan, et al.
Pubblicazione: (2025) -
NAACL2025 Tutorial: Adaptation of Large Language Models
di: Ke, Zixuan, et al.
Pubblicazione: (2025)