Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
Fuente:
arXiv
Guardado en:
| Autores principales: | Moell, Birger, Aronsson, Fredrik Sand, Akbar, Sanian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
por: Moell, Birger, et al.
Publicado: (2025)
por: Moell, Birger, et al.
Publicado: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
por: Moell, Birger, et al.
Publicado: (2025)
por: Moell, Birger, et al.
Publicado: (2025)
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
por: Shmidman, Shaltiel, et al.
Publicado: (2025)
por: Shmidman, Shaltiel, et al.
Publicado: (2025)
Evaluating Large Language Models with Human Feedback: Establishing a Swedish Benchmark
por: Moell, Birger
Publicado: (2024)
por: Moell, Birger
Publicado: (2024)
Comparing the Efficacy of GPT-4 and Chat-GPT in Mental Health Care: A Blind Assessment of Large Language Models for Psychological Support
por: Moell, Birger
Publicado: (2024)
por: Moell, Birger
Publicado: (2024)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
por: DeepSeek-AI, et al.
Publicado: (2025)
por: DeepSeek-AI, et al.
Publicado: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
por: Marjanović, Sara Vera, et al.
Publicado: (2025)
por: Marjanović, Sara Vera, et al.
Publicado: (2025)
A Comparison of DeepSeek and Other LLMs
por: Gao, Tianchen, et al.
Publicado: (2025)
por: Gao, Tianchen, et al.
Publicado: (2025)
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
por: Hu, Yinghao, et al.
Publicado: (2025)
por: Hu, Yinghao, et al.
Publicado: (2025)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
por: Larionov, Daniil, et al.
Publicado: (2025)
por: Larionov, Daniil, et al.
Publicado: (2025)
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance
por: Moell, Birger, et al.
Publicado: (2025)
por: Moell, Birger, et al.
Publicado: (2025)
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
Explainable Sentiment Analysis with DeepSeek-R1: Performance, Efficiency, and Few-Shot Learning
por: Huang, Donghao, et al.
Publicado: (2025)
por: Huang, Donghao, et al.
Publicado: (2025)
Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
por: Li, Xuying, et al.
Publicado: (2025)
por: Li, Xuying, et al.
Publicado: (2025)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
por: Zhang, Jue, et al.
Publicado: (2025)
por: Zhang, Jue, et al.
Publicado: (2025)
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
por: Zhang, Chong, et al.
Publicado: (2025)
por: Zhang, Chong, et al.
Publicado: (2025)
Mixture of Tunable Experts -- Behavior Modification of DeepSeek-R1 at Inference Time
por: Dahlke, Robert, et al.
Publicado: (2025)
por: Dahlke, Robert, et al.
Publicado: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
por: Naseh, Ali, et al.
Publicado: (2025)
por: Naseh, Ali, et al.
Publicado: (2025)
DeepSeek-V3 Technical Report
por: DeepSeek-AI, et al.
Publicado: (2024)
por: DeepSeek-AI, et al.
Publicado: (2024)
Safety Evaluation of DeepSeek Models in Chinese Contexts
por: Zhang, Wenjing, et al.
Publicado: (2025)
por: Zhang, Wenjing, et al.
Publicado: (2025)
LLMs in Disease Diagnosis: A Comparative Study of DeepSeek-R1 and O3 Mini Across Chronic Health Conditions
por: Gupta, Gaurav Kumar, et al.
Publicado: (2025)
por: Gupta, Gaurav Kumar, et al.
Publicado: (2025)
Artificial Humans
por: Moell, Birger
Publicado: (2025)
por: Moell, Birger
Publicado: (2025)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
por: Zhang, Wenjing, et al.
Publicado: (2025)
por: Zhang, Wenjing, et al.
Publicado: (2025)
An evaluation of DeepSeek Models in Biomedical Natural Language Processing
por: Zhan, Zaifu, et al.
Publicado: (2025)
por: Zhan, Zaifu, et al.
Publicado: (2025)
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish
por: Cumbal, Ronald, et al.
Publicado: (2024)
por: Cumbal, Ronald, et al.
Publicado: (2024)
Emotion-Aware Embedding Fusion in LLMs (Flan-T5, LLAMA 2, DeepSeek-R1, and ChatGPT 4) for Intelligent Response Generation
por: Rasool, Abdur, et al.
Publicado: (2024)
por: Rasool, Abdur, et al.
Publicado: (2024)
Comparative Analysis of OpenAI GPT-4o and DeepSeek R1 for Scientific Text Categorization Using Prompt Engineering
por: Maiti, Aniruddha, et al.
Publicado: (2025)
por: Maiti, Aniruddha, et al.
Publicado: (2025)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
por: Ji, Tao, et al.
Publicado: (2025)
por: Ji, Tao, et al.
Publicado: (2025)
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
por: DeepSeek-AI, et al.
Publicado: (2025)
por: DeepSeek-AI, et al.
Publicado: (2025)
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
por: Ren, Z. Z., et al.
Publicado: (2025)
por: Ren, Z. Z., et al.
Publicado: (2025)
Reasoning and the Trusting Behavior of DeepSeek and GPT: An Experiment Revealing Hidden Fault Lines in Large Language Models
por: Li, Rubing, et al.
Publicado: (2025)
por: Li, Rubing, et al.
Publicado: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
por: Xu, Pusheng, et al.
Publicado: (2025)
por: Xu, Pusheng, et al.
Publicado: (2025)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
por: Parmar, Manojkumar, et al.
Publicado: (2025)
por: Parmar, Manojkumar, et al.
Publicado: (2025)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
por: Liang, Yunhao, et al.
Publicado: (2026)
por: Liang, Yunhao, et al.
Publicado: (2026)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
por: Sands, Brendan, et al.
Publicado: (2025)
por: Sands, Brendan, et al.
Publicado: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
por: DeepSeek-AI, et al.
Publicado: (2024)
por: DeepSeek-AI, et al.
Publicado: (2024)
Analysis of LLM Bias (Chinese Propaganda & Anti-US Sentiment) in DeepSeek-R1 vs. ChatGPT o3-mini-high
por: Huang, PeiHsuan, et al.
Publicado: (2025)
por: Huang, PeiHsuan, et al.
Publicado: (2025)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
por: Ma, Boyi, et al.
Publicado: (2025)
por: Ma, Boyi, et al.
Publicado: (2025)
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
por: Kuo, Martin, et al.
Publicado: (2025)
por: Kuo, Martin, et al.
Publicado: (2025)
Ejemplares similares
-
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
por: Moell, Birger, et al.
Publicado: (2025) -
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
por: Moell, Birger, et al.
Publicado: (2025) -
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
por: Shmidman, Shaltiel, et al.
Publicado: (2025) -
Evaluating Large Language Models with Human Feedback: Establishing a Swedish Benchmark
por: Moell, Birger
Publicado: (2024) -
Comparing the Efficacy of GPT-4 and Chat-GPT in Mental Health Care: A Blind Assessment of Large Language Models for Psychological Support
por: Moell, Birger
Publicado: (2024)