MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Minpeng, Luo, Wei, Li, Chengxi, Wu, Jing, Fan, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Step-level Value Preference Optimization for Mathematical Reasoning
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
AlphaMath Almost Zero: Process Supervision without Process
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
Markov Chain of Thought for Efficient Mathematical Reasoning
by: Yang, Wen, et al.
Published: (2024)
by: Yang, Wen, et al.
Published: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
by: Zhang, Boning, et al.
Published: (2024)
by: Zhang, Boning, et al.
Published: (2024)
LLMs Can Achieve High-quality Simultaneous Machine Translation as Efficiently as Offline
by: Fu, Biao, et al.
Published: (2025)
by: Fu, Biao, et al.
Published: (2025)
Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
by: Fu, Biao, et al.
Published: (2025)
by: Fu, Biao, et al.
Published: (2025)
ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
by: Cheng, Qianjia, et al.
Published: (2026)
by: Cheng, Qianjia, et al.
Published: (2026)
A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text
by: Mikros, George, et al.
Published: (2026)
by: Mikros, George, et al.
Published: (2026)
RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
by: Duarte, André V., et al.
Published: (2025)
by: Duarte, André V., et al.
Published: (2025)
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
by: Luo, Linhao, et al.
Published: (2023)
by: Luo, Linhao, et al.
Published: (2023)
DRS: Deep Question Reformulation With Structured Output
by: Li, Zhecheng, et al.
Published: (2024)
by: Li, Zhecheng, et al.
Published: (2024)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Learning a Structural Causal Model for Intuition Reasoning in Conversation
by: Chen, Hang, et al.
Published: (2023)
by: Chen, Hang, et al.
Published: (2023)
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
by: Chen, Xinjie, et al.
Published: (2025)
by: Chen, Xinjie, et al.
Published: (2025)
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
by: Lu, Qiuyang, et al.
Published: (2025)
by: Lu, Qiuyang, et al.
Published: (2025)
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
by: Huang, You-Liang, et al.
Published: (2026)
by: Huang, You-Liang, et al.
Published: (2026)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
by: Jing, Huihao, et al.
Published: (2025)
by: Jing, Huihao, et al.
Published: (2025)
CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction
by: Li, Junlong, et al.
Published: (2025)
by: Li, Junlong, et al.
Published: (2025)
Data Augmented Pipeline for Legal Information Extraction and Reasoning
by: Phuong, Nguyen Minh, et al.
Published: (2026)
by: Phuong, Nguyen Minh, et al.
Published: (2026)
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling
by: Zhang, Kai, et al.
Published: (2026)
by: Zhang, Kai, et al.
Published: (2026)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
by: Chang, Shuochen, et al.
Published: (2026)
by: Chang, Shuochen, et al.
Published: (2026)
What Is Missing: Interpretable Ratings for Large Language Model Outputs
by: Stranges, Nicholas, et al.
Published: (2026)
by: Stranges, Nicholas, et al.
Published: (2026)
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
by: Wang, Chen, et al.
Published: (2023)
by: Wang, Chen, et al.
Published: (2023)
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
by: Chen, Hailin, et al.
Published: (2024)
by: Chen, Hailin, et al.
Published: (2024)
PARM: Pipeline-Adapted Reward Model
by: Fan, Xingyu, et al.
Published: (2026)
by: Fan, Xingyu, et al.
Published: (2026)
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression
by: Li, Yuankai, et al.
Published: (2024)
by: Li, Yuankai, et al.
Published: (2024)
A Course Shared Task on Evaluating LLM Output for Clinical Questions
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
by: Sun, Chung-En, et al.
Published: (2025)
by: Sun, Chung-En, et al.
Published: (2025)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
A New Pipeline For Generating Instruction Dataset via RAG and Self Fine-Tuning
by: Song, Chih-Wei, et al.
Published: (2024)
by: Song, Chih-Wei, et al.
Published: (2024)
G2: Guided Generation for Enhanced Output Diversity in LLMs
by: Ruan, Zhiwen, et al.
Published: (2025)
by: Ruan, Zhiwen, et al.
Published: (2025)
How Interpretable are Reasoning Explanations from Prompting Large Language Models?
by: Yeo, Wei Jie, et al.
Published: (2024)
by: Yeo, Wei Jie, et al.
Published: (2024)
OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models
by: Yue, Yuanhao, et al.
Published: (2026)
by: Yue, Yuanhao, et al.
Published: (2026)
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Similar Items
-
Step-level Value Preference Optimization for Mathematical Reasoning
by: Chen, Guoxin, et al.
Published: (2024) -
AlphaMath Almost Zero: Process Supervision without Process
by: Chen, Guoxin, et al.
Published: (2024) -
Markov Chain of Thought for Efficient Mathematical Reasoning
by: Yang, Wen, et al.
Published: (2024) -
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
by: Zhang, Boning, et al.
Published: (2024) -
LLMs Can Achieve High-quality Simultaneous Machine Translation as Efficiently as Offline
by: Fu, Biao, et al.
Published: (2025)