DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Pusheng, Wu, Yue, Jin, Kai, Chen, Xiaolan, He, Mingguang, Shi, Danli |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
by: Dong, Ben, et al.
Published: (2025)
by: Dong, Ben, et al.
Published: (2025)
Memory Analysis on the Training Course of DeepSeek Models
by: Zhang, Ping, et al.
Published: (2025)
by: Zhang, Ping, et al.
Published: (2025)
Visual Question Answering in Ophthalmology: A Progressive and Practical Perspective
by: Chen, Xiaolan, et al.
Published: (2024)
by: Chen, Xiaolan, et al.
Published: (2024)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
by: Pfister, Rolf, et al.
Published: (2025)
by: Pfister, Rolf, et al.
Published: (2025)
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
by: Kuo, Martin, et al.
Published: (2025)
by: Kuo, Martin, et al.
Published: (2025)
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
by: Hu, Yinghao, et al.
Published: (2025)
by: Hu, Yinghao, et al.
Published: (2025)
o3-mini vs DeepSeek-R1: Which One is Safer?
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
Outperforming Multiserver SRPT at All Loads
by: Grosof, Izzy, et al.
Published: (2025)
by: Grosof, Izzy, et al.
Published: (2025)
Comparative Analysis of OpenAI GPT-4o and DeepSeek R1 for Scientific Text Categorization Using Prompt Engineering
by: Maiti, Aniruddha, et al.
Published: (2025)
by: Maiti, Aniruddha, et al.
Published: (2025)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
by: Kumar, Deepak, et al.
Published: (2025)
by: Kumar, Deepak, et al.
Published: (2025)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
by: Larionov, Daniil, et al.
Published: (2025)
by: Larionov, Daniil, et al.
Published: (2025)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
by: Sands, Brendan, et al.
Published: (2025)
by: Sands, Brendan, et al.
Published: (2025)
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3
by: Sadik, Ahmed R., et al.
Published: (2025)
by: Sadik, Ahmed R., et al.
Published: (2025)
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Accuracy of ChatGPT , Gemini, Claude and DeepSeek in Carbohydrate Counting
by: Luca Zagaroli, et al.
Published: (2026)
by: Luca Zagaroli, et al.
Published: (2026)
EyeDiff: text-to-image diffusion model improves rare eye disease diagnosis
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study
by: Srinivasan, Sahana, et al.
Published: (2025)
by: Srinivasan, Sahana, et al.
Published: (2025)
Outperforming Dijkstra on Sparse Graphs: The Lightning Network Use Case
by: Valko, Danila, et al.
Published: (2025)
by: Valko, Danila, et al.
Published: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
Da artificação do sagrado nos museus: entre o teatro e a sacralidade
by: Bruno Brulon
Published: (2013)
by: Bruno Brulon
Published: (2013)
Analysis of LLM Bias (Chinese Propaganda & Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high
by: Huang, PeiHsuan, et al.
Published: (2025)
by: Huang, PeiHsuan, et al.
Published: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
by: DeepSeek-AI, et al.
Published: (2025)
by: DeepSeek-AI, et al.
Published: (2025)
Cleaning up the Mess: Re-Evaluating the Real-System Modeling Accuracy of Ramulator 2.0
by: Bostanci, F. Nisa, et al.
Published: (2025)
by: Bostanci, F. Nisa, et al.
Published: (2025)
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
by: Shmidman, Shaltiel, et al.
Published: (2025)
by: Shmidman, Shaltiel, et al.
Published: (2025)
Integrating ytopt and libEnsemble to Autotune OpenMC
by: Wu, Xingfu, et al.
Published: (2024)
by: Wu, Xingfu, et al.
Published: (2024)
ppOpen-AT: A Directive-base Auto-tuning Language
by: Katagiri, Takahiro
Published: (2024)
by: Katagiri, Takahiro
Published: (2024)
A Microbenchmark Framework for Performance Evaluation of OpenMP Target Offloading
by: Atif, Mohammad, et al.
Published: (2025)
by: Atif, Mohammad, et al.
Published: (2025)
Who Knows Anatomy Best? A Comparative Study of ChatGPT ‐4o, DeepSeek , Gemini, and Claude
by: Melek Tassoker
Published: (2025)
by: Melek Tassoker
Published: (2025)
EyeGPT: Ophthalmic Assistant with Large Language Models
by: Chen, Xiaolan, et al.
Published: (2024)
by: Chen, Xiaolan, et al.
Published: (2024)
A Priori Loop Nest Normalization: Automatic Loop Scheduling in Complex Applications
by: Trümper, Lukas, et al.
Published: (2024)
by: Trümper, Lukas, et al.
Published: (2024)
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Industry 4.0 Connectors -- A Performance Experiment with Modbus/TCP
by: Nikolajew, Christian, et al.
Published: (2024)
by: Nikolajew, Christian, et al.
Published: (2024)
A Controlled Study of Memory Hierarchy Transitions in Quantum Circuit Simulation on Apple M4 Pro Unified Memory Architecture
by: Pratipat, Gyan
Published: (2026)
by: Pratipat, Gyan
Published: (2026)
PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
Assessing the Performance of OpenTitan as Cryptographic Accelerator in Secure Open-Hardware System-on-Chips
by: Parisi, Emanuele, et al.
Published: (2024)
by: Parisi, Emanuele, et al.
Published: (2024)
Accuracy of Generative AI Chatbots in Answering Plastic Surgery Examination Questions: A Comparative Evaluation of ChatGPT‐4o, Gemini Advanced, and DeepSeek‐R1
by: Jiaxian Zhang, et al.
Published: (2026)
by: Jiaxian Zhang, et al.
Published: (2026)
Memshare: Memory Sharing for Multicore Computation in R with an Application to Feature Selection by Mutual Information using PDE
by: Thrun, Michael C., et al.
Published: (2025)
by: Thrun, Michael C., et al.
Published: (2025)
Efficient Transpilation of OpenQASM 3.0 Dynamic Circuits to CUDA-Q: Performance and Expressiveness Advantages
by: Kulkarni, Vinooth, et al.
Published: (2026)
by: Kulkarni, Vinooth, et al.
Published: (2026)
Similar Items
-
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
by: Dong, Ben, et al.
Published: (2025) -
Memory Analysis on the Training Course of DeepSeek Models
by: Zhang, Ping, et al.
Published: (2025) -
Visual Question Answering in Ophthalmology: A Progressive and Practical Perspective
by: Chen, Xiaolan, et al.
Published: (2024) -
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
by: Pfister, Rolf, et al.
Published: (2025) -
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
by: Kuo, Martin, et al.
Published: (2025)