Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yu, Yin, Huifeng, Zeng, Bo, Wang, Hao, Shi, Tianqi, Lyu, Chenyang, Wang, Longyue, Luo, Weihua, Zhang, Kaifu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929603568730112
author Zhao, Yu
Yin, Huifeng
Zeng, Bo
Wang, Hao
Shi, Tianqi
Lyu, Chenyang
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
author_facet Zhao, Yu
Yin, Huifeng
Zeng, Bo
Wang, Hao
Shi, Tianqi
Lyu, Chenyang
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
contents Currently OpenAI o1 sparks a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mathematics, physics, and coding -- which are well-suited for reinforcement learning (RL) -- but also places greater emphasis on open-ended resolutions. We aim to address the question: ''Can the o1 model effectively generalize to broader domains where clear standards are absent and rewards are challenging to quantify?'' Marco-o1 is powered by Chain-of-Thought (CoT) fine-tuning, Monte Carlo Tree Search (MCTS), reflection mechanisms, and innovative reasoning strategies -- optimized for complex real-world problem-solving tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14405
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
Zhao, Yu
Yin, Huifeng
Zeng, Bo
Wang, Hao
Shi, Tianqi
Lyu, Chenyang
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
Computation and Language
Currently OpenAI o1 sparks a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mathematics, physics, and coding -- which are well-suited for reinforcement learning (RL) -- but also places greater emphasis on open-ended resolutions. We aim to address the question: ''Can the o1 model effectively generalize to broader domains where clear standards are absent and rewards are challenging to quantify?'' Marco-o1 is powered by Chain-of-Thought (CoT) fine-tuning, Monte Carlo Tree Search (MCTS), reflection mechanisms, and innovative reasoning strategies -- optimized for complex real-world problem-solving tasks.
title Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
topic Computation and Language
url https://arxiv.org/abs/2411.14405