O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zhen, Zou, Haoyang, Li, Xuefeng, Liu, Yixiu, Zheng, Yuxiang, Chern, Ethan, Xia, Shijie, Qin, Yiwei, Yuan, Weizhe, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
O1 Replication Journey: A Strategic Progress Report -- Part 1
by: Qin, Yiwei, et al.
Published: (2024)
by: Qin, Yiwei, et al.
Published: (2024)
O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
by: Huang, Zhongzhen, et al.
Published: (2025)
by: Huang, Zhongzhen, et al.
Published: (2025)
DIVE: Diversified Iterative Self-Improvement
by: Qin, Yiwei, et al.
Published: (2025)
by: Qin, Yiwei, et al.
Published: (2025)
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
by: Xia, Shijie, et al.
Published: (2025)
by: Xia, Shijie, et al.
Published: (2025)
ToRL: Scaling Tool-Integrated RL
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Reformatted Alignment
by: Fan, Run-Ze, et al.
Published: (2024)
by: Fan, Run-Ze, et al.
Published: (2024)
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024)
by: Wang, Binjie, et al.
Published: (2024)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
SAFETY-J: Evaluating Safety with Critique
by: Liu, Yixiu, et al.
Published: (2024)
by: Liu, Yixiu, et al.
Published: (2024)
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
by: He, Yanheng, et al.
Published: (2024)
by: He, Yanheng, et al.
Published: (2024)
LIMO: Less is More for Reasoning
by: Ye, Yixin, et al.
Published: (2025)
by: Ye, Yixin, et al.
Published: (2025)
A Bitter Lesson for Data Filtering
by: Mohri, Christopher, et al.
Published: (2026)
by: Mohri, Christopher, et al.
Published: (2026)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
by: Chern, Ethan, et al.
Published: (2024)
by: Chern, Ethan, et al.
Published: (2024)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
Code for: The Second Dirichlet Eigenvalue is Simple on Every Non-equilateral Triangle, Part II
by: Endo, Ryoki, et al.
Published: (2025)
by: Endo, Ryoki, et al.
Published: (2025)
Progress or Regress? Self-Improvement Reversal in Post-training
by: Wu, Ting, et al.
Published: (2024)
by: Wu, Ting, et al.
Published: (2024)
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
by: Wu, Minghao, et al.
Published: (2025)
by: Wu, Minghao, et al.
Published: (2025)
Data Darwinism Part II: DataEvolve -- AI can Autonomously Evolve Pretraining Data Curation
by: Mi, Tiantian, et al.
Published: (2026)
by: Mi, Tiantian, et al.
Published: (2026)
Show preview
by: Anonimo
Published: (2009)
by: Anonimo
Published: (2009)
News and previews
Published: (2004)
Published: (2004)
The Second Dirichlet Eigenvalue is Simple on Every Non-equilateral Triangle, Part II: Nearly Equilateral Triangles
by: Endo, Ryoki, et al.
Published: (2023)
by: Endo, Ryoki, et al.
Published: (2023)
Alignment for Honesty
by: Yang, Yuqing, et al.
Published: (2023)
by: Yang, Yuqing, et al.
Published: (2023)
Evaluating Mathematical Reasoning Beyond Accuracy
by: Xia, Shijie, et al.
Published: (2024)
by: Xia, Shijie, et al.
Published: (2024)
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Thinking with Generated Images
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models
by: Xiaoyu, Wang, et al.
Published: (2026)
by: Xiaoyu, Wang, et al.
Published: (2026)
The Brain's Bitter Lesson: Scaling Speech Decoding With Self-Supervised Learning
by: Jayalath, Dulhan, et al.
Published: (2024)
by: Jayalath, Dulhan, et al.
Published: (2024)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
LLMCRIT: Teaching Large Language Models to Use Criteria
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
Letter: The Simple Endoscopic Score for Predicting Crohn's Disease Progression—Not as Simple as It Sounds
by: Ziheng Calvin Xu, et al.
Published: (2025)
by: Ziheng Calvin Xu, et al.
Published: (2025)
Learning the Bitter Lesson: Empirical Evidence from 20 Years of CVPR Proceedings
by: Yousefi, Mojtaba, et al.
Published: (2024)
by: Yousefi, Mojtaba, et al.
Published: (2024)
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
by: Lu, Qingyu, et al.
Published: (2026)
by: Lu, Qingyu, et al.
Published: (2026)
Adversarial Score identity Distillation: Rapidly Surpassing the Teacher in One Step
by: Zhou, Mingyuan, et al.
Published: (2024)
by: Zhou, Mingyuan, et al.
Published: (2024)
o1-Coder: an o1 Replication for Coding
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
A note on knot Floer homology of satellite knots with (1,1)-patterns
by: Shen, Weizhe
Published: (2022)
by: Shen, Weizhe
Published: (2022)
Study of Sensory Evaluation of Bittering Agents for Establishing an Evaluation System of Bitterness
by: Kaho Watanabe, et al.
Published: (2026)
by: Kaho Watanabe, et al.
Published: (2026)
A Sieve Method for Generating 1+1 Prime Pairs and Goldbach Function
by: Chern, Geeng-Chuan
Published: (2026)
by: Chern, Geeng-Chuan
Published: (2026)
Bitter-Sweet Democracy?
Published: (2024)
Published: (2024)
Similar Items
-
O1 Replication Journey: A Strategic Progress Report -- Part 1
by: Qin, Yiwei, et al.
Published: (2024) -
O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
by: Huang, Zhongzhen, et al.
Published: (2025) -
DIVE: Diversified Iterative Self-Improvement
by: Qin, Yiwei, et al.
Published: (2025) -
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
by: Xia, Shijie, et al.
Published: (2025) -
ToRL: Scaling Tool-Integrated RL
by: Li, Xuefeng, et al.
Published: (2025)