Mixture of Reasonings: Teach Large Language Models to Reason with Adaptive Strategies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Tao, Hu, Xavier, Fan, Wenyan, Zhang, Shengyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908431209725952
author Xiong, Tao
Hu, Xavier
Fan, Wenyan
Zhang, Shengyu
author_facet Xiong, Tao
Hu, Xavier
Fan, Wenyan
Zhang, Shengyu
contents Large language models (LLMs) excel in complex tasks through advanced prompting techniques like Chain-of-Thought (CoT) and Tree-of-Thought (ToT), but their reliance on manually crafted, task-specific prompts limits adaptability and efficiency. We introduce Mixture of Reasoning (MoR), a training framework that embeds diverse reasoning strategies into LLMs for autonomous, task-adaptive reasoning without external prompt engineering. MoR has two phases: Thought Generation, creating reasoning chain templates with models like GPT-4o, and SFT Dataset Construction, pairing templates with benchmark datasets for supervised fine-tuning. Our experiments show that MoR significantly enhances performance, with MoR150 achieving 0.730 (2.2% improvement) using CoT prompting and 0.734 (13.5% improvement) compared to baselines. MoR eliminates the need for task-specific prompts, offering a generalizable solution for robust reasoning across diverse tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture of Reasonings: Teach Large Language Models to Reason with Adaptive Strategies
Xiong, Tao
Hu, Xavier
Fan, Wenyan
Zhang, Shengyu
Computation and Language
Artificial Intelligence
Large language models (LLMs) excel in complex tasks through advanced prompting techniques like Chain-of-Thought (CoT) and Tree-of-Thought (ToT), but their reliance on manually crafted, task-specific prompts limits adaptability and efficiency. We introduce Mixture of Reasoning (MoR), a training framework that embeds diverse reasoning strategies into LLMs for autonomous, task-adaptive reasoning without external prompt engineering. MoR has two phases: Thought Generation, creating reasoning chain templates with models like GPT-4o, and SFT Dataset Construction, pairing templates with benchmark datasets for supervised fine-tuning. Our experiments show that MoR significantly enhances performance, with MoR150 achieving 0.730 (2.2% improvement) using CoT prompting and 0.734 (13.5% improvement) compared to baselines. MoR eliminates the need for task-specific prompts, offering a generalizable solution for robust reasoning across diverse tasks.
title Mixture of Reasonings: Teach Large Language Models to Reason with Adaptive Strategies
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.00606