PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Xuekai, Qi, Biqing, Zhang, Kaiyan, Long, Xinwei, Lin, Zhouhan, Zhou, Bowen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Critical Data Size of Language Models from a Grokking Perspective
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
by: Long, Xinwei, et al.
Published: (2025)
by: Long, Xinwei, et al.
Published: (2025)
Towards Building Specialized Generalist AI with System 1 and System 2 Fusion
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
by: Hua, Ermo, et al.
Published: (2024)
by: Hua, Ermo, et al.
Published: (2024)
Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
How to Synthesize Text Data without Model Collapse?
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
Pole-centric Descriptors for Robust Robot Localization: Evaluation under Pole-at-Distance (PaD) Observations using the Small Pole Landmark (SPL) Dataset
by: Xie, Wuhao, et al.
Published: (2025)
by: Xie, Wuhao, et al.
Published: (2025)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
by: Hua, Ermo, et al.
Published: (2024)
by: Hua, Ermo, et al.
Published: (2024)
Evolution of Thought: Diverse and High-Quality Reasoning via Multi-Objective Optimization
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
by: Chen, Lin, et al.
Published: (2026)
by: Chen, Lin, et al.
Published: (2026)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
by: Hościłowicz, Jakub, et al.
Published: (2023)
by: Hościłowicz, Jakub, et al.
Published: (2023)
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Improve Vision Language Model Chain-of-thought Reasoning
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?
by: Zhou, Zhanke, et al.
Published: (2024)
by: Zhou, Zhanke, et al.
Published: (2024)
Early Stopping Chain-of-thoughts in Large Language Models
by: Mao, Minjia, et al.
Published: (2025)
by: Mao, Minjia, et al.
Published: (2025)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
TTRL: Test-Time Reinforcement Learning
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
MeLo: Low-rank Adaptation is Better than Fine-tuning for Medical Image Diagnosis
by: Zhu, Yitao, et al.
Published: (2023)
by: Zhu, Yitao, et al.
Published: (2023)
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
by: Huang, Chenxi, et al.
Published: (2025)
by: Huang, Chenxi, et al.
Published: (2025)
Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models
by: Lv, Xingtai, et al.
Published: (2025)
by: Lv, Xingtai, et al.
Published: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
by: Fu, Tianhao, et al.
Published: (2025)
by: Fu, Tianhao, et al.
Published: (2025)
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
by: Ning, Xuefei, et al.
Published: (2024)
by: Ning, Xuefei, et al.
Published: (2024)
SMR: State Memory Replay for Long Sequence Modeling
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
ReviewRL: Towards Automated Scientific Review with RL
by: Zeng, Sihang, et al.
Published: (2025)
by: Zeng, Sihang, et al.
Published: (2025)
Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought
by: Lee, Jooyoung, et al.
Published: (2024)
by: Lee, Jooyoung, et al.
Published: (2024)
Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNets
by: Xue, Bo, et al.
Published: (2026)
by: Xue, Bo, et al.
Published: (2026)
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
by: Zhao, Jian, et al.
Published: (2025)
by: Zhao, Jian, et al.
Published: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Generative Multi-Modal Knowledge Retrieval with Large Language Models
by: Long, Xinwei, et al.
Published: (2024)
by: Long, Xinwei, et al.
Published: (2024)
SR-CIS: Self-Reflective Incremental System with Decoupled Memory and Reasoning
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
A Chain-of-thought Reasoning Breast Ultrasound Dataset Covering All Histopathology Categories
by: Yu, Haojun, et al.
Published: (2025)
by: Yu, Haojun, et al.
Published: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
by: Zhu, Xuekai, et al.
Published: (2025)
by: Zhu, Xuekai, et al.
Published: (2025)
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
by: Sprague, Zayne, et al.
Published: (2023)
by: Sprague, Zayne, et al.
Published: (2023)
MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
by: Dutta, Aritra, et al.
Published: (2026)
by: Dutta, Aritra, et al.
Published: (2026)
Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative Watermarking
by: Ma, Zhiyuan, et al.
Published: (2024)
by: Ma, Zhiyuan, et al.
Published: (2024)
Neural Residual Diffusion Models for Deep Scalable Vision Generation
by: Ma, Zhiyuan, et al.
Published: (2024)
by: Ma, Zhiyuan, et al.
Published: (2024)
MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning
by: Li, Nianqi, et al.
Published: (2024)
by: Li, Nianqi, et al.
Published: (2024)
Similar Items
-
Critical Data Size of Language Models from a Grokking Perspective
by: Zhu, Xuekai, et al.
Published: (2024) -
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
by: Long, Xinwei, et al.
Published: (2025) -
Towards Building Specialized Generalist AI with System 1 and System 2 Fusion
by: Zhang, Kaiyan, et al.
Published: (2024) -
CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
by: Zhang, Kaiyan, et al.
Published: (2024) -
Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
by: Hua, Ermo, et al.
Published: (2024)