OpenAI-o1 AB Testing: Does the o1 model really do good reasoning in math problem solving?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Leo, Luo, Ye, Pan, Tingyou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenAI o1 System Card
by: OpenAI, et al.
Published: (2024)
by: OpenAI, et al.
Published: (2024)
When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1
by: McCoy, R. Thomas, et al.
Published: (2024)
by: McCoy, R. Thomas, et al.
Published: (2024)
On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability
by: Wang, Kevin, et al.
Published: (2024)
by: Wang, Kevin, et al.
Published: (2024)
Can OpenAI o1 outperform humans in higher-order cognitive thinking?
by: Latif, Ehsan, et al.
Published: (2024)
by: Latif, Ehsan, et al.
Published: (2024)
A Systematic Assessment of OpenAI o1-Preview for Higher Order Thinking in Education
by: Latif, Ehsan, et al.
Published: (2024)
by: Latif, Ehsan, et al.
Published: (2024)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
by: Pfister, Rolf, et al.
Published: (2025)
by: Pfister, Rolf, et al.
Published: (2025)
Testing GPT-4-o1-preview on math and science problems: A follow-up study
by: Davis, Ernest
Published: (2024)
by: Davis, Ernest
Published: (2024)
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam
by: de Winter, Joost, et al.
Published: (2024)
by: de Winter, Joost, et al.
Published: (2024)
OpenAI GPT-5 System Card
by: Singh, Aaditya, et al.
Published: (2025)
by: Singh, Aaditya, et al.
Published: (2025)
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
by: Han, Yifu, et al.
Published: (2025)
by: Han, Yifu, et al.
Published: (2025)
Privacy and Security Threat for OpenAI GPTs
by: Wenying, Wei, et al.
Published: (2025)
by: Wenying, Wei, et al.
Published: (2025)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
by: Valmeekam, Karthik, et al.
Published: (2024)
by: Valmeekam, Karthik, et al.
Published: (2024)
The Phenomenology of Machine: A Comprehensive Analysis of the Sentience of the OpenAI-o1 Model Integrating Functionalism, Consciousness Theories, Active Inference, and AI Architectures
by: Hoyle, Victoria Violet
Published: (2024)
by: Hoyle, Victoria Violet
Published: (2024)
Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study
by: Srinivasan, Sahana, et al.
Published: (2025)
by: Srinivasan, Sahana, et al.
Published: (2025)
How well do LLMs reason over tabular data, really?
by: Wolff, Cornelius, et al.
Published: (2025)
by: Wolff, Cornelius, et al.
Published: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024)
by: Sprague, Zayne, et al.
Published: (2024)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
by: Xu, Pusheng, et al.
Published: (2025)
by: Xu, Pusheng, et al.
Published: (2025)
An Empirical Study of OpenAI API Discussions on Stack Overflow
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
On stress-related behavior patterns in OpenAI's o3-mini model: analysis and mitigations
by: Levchenko, Anastasia, et al.
Published: (2025)
by: Levchenko, Anastasia, et al.
Published: (2025)
Give me a hint: Can LLMs take a hint to solve math problems?
by: Agrawal, Vansh, et al.
Published: (2024)
by: Agrawal, Vansh, et al.
Published: (2024)
Comparative Analysis of OpenAI GPT-4o and DeepSeek R1 for Scientific Text Categorization Using Prompt Engineering
by: Maiti, Aniruddha, et al.
Published: (2025)
by: Maiti, Aniruddha, et al.
Published: (2025)
OpenAI's Approach to External Red Teaming for AI Models and Systems
by: Ahmad, Lama, et al.
Published: (2025)
by: Ahmad, Lama, et al.
Published: (2025)
Benchmarking Floworks against OpenAI & Anthropic: A Novel Framework for Enhanced LLM Function Calling
by: Bhan, Nirav, et al.
Published: (2024)
by: Bhan, Nirav, et al.
Published: (2024)
A Case Study of Web App Coding with OpenAI Reasoning Models
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems
by: Davis, Ernest, et al.
Published: (2023)
by: Davis, Ernest, et al.
Published: (2023)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
FinAI Data Assistant: LLM-based Financial Database Query Processing with the OpenAI Function Calling API
by: Kim, Juhyeong, et al.
Published: (2025)
by: Kim, Juhyeong, et al.
Published: (2025)
MTUncertainty: Assessing the Need for Post-editing of Machine Translation Outputs by Fine-tuning OpenAI LLMs
by: Gladkoff, Serge, et al.
Published: (2023)
by: Gladkoff, Serge, et al.
Published: (2023)
Fast Analysis of the OpenAI O1-Preview Model in Solving Random K-SAT Problem: Does the LLM Solve the Problem Itself or Call an External SAT Solver?
by: Marino, Raffaele
Published: (2024)
by: Marino, Raffaele
Published: (2024)
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
by: Shendabadi, Ali, et al.
Published: (2026)
by: Shendabadi, Ali, et al.
Published: (2026)
In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
by: Durner, Nils
Published: (2025)
by: Durner, Nils
Published: (2025)
OpenAI's GPT-OSS-20B Model and Safety Alignment Issues in a Low-Resource Language
by: Inuwa-Dutse, Isa
Published: (2025)
by: Inuwa-Dutse, Isa
Published: (2025)
A Large-Scale Empirical Analysis of Custom GPTs' Vulnerabilities in the OpenAI Ecosystem
by: Ogundoyin, Sunday Oyinlola, et al.
Published: (2025)
by: Ogundoyin, Sunday Oyinlola, et al.
Published: (2025)
Proof-of-TBI -- Fine-Tuned Vision Language Model Consortium and OpenAI-o3 Reasoning LLM-Based Medical Diagnosis Support System for Mild Traumatic Brain Injury (TBI) Prediction
by: Gore, Ross, et al.
Published: (2025)
by: Gore, Ross, et al.
Published: (2025)
HDDLGym: A Tool for Studying Multi-Agent Hierarchical Problems Defined in HDDL with OpenAI Gym
by: La, Ngoc, et al.
Published: (2025)
by: La, Ngoc, et al.
Published: (2025)
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
by: Irugalbandara, Chandra, et al.
Published: (2023)
by: Irugalbandara, Chandra, et al.
Published: (2023)
Standardization of Psychiatric Diagnoses -- Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System
by: Bandara, Eranga, et al.
Published: (2025)
by: Bandara, Eranga, et al.
Published: (2025)
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
by: Coggins, Sam, et al.
Published: (2025)
by: Coggins, Sam, et al.
Published: (2025)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
by: Freeman, Joshua, et al.
Published: (2024)
by: Freeman, Joshua, et al.
Published: (2024)
Similar Items
-
OpenAI o1 System Card
by: OpenAI, et al.
Published: (2024) -
When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1
by: McCoy, R. Thomas, et al.
Published: (2024) -
On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability
by: Wang, Kevin, et al.
Published: (2024) -
Can OpenAI o1 outperform humans in higher-order cognitive thinking?
by: Latif, Ehsan, et al.
Published: (2024) -
A Systematic Assessment of OpenAI o1-Preview for Higher Order Thinking in Education
by: Latif, Ehsan, et al.
Published: (2024)