Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH
Fuente:
arXiv
Salvato in:
| Autore principale: | Evstafev, Evgenii |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Token-by-Token Regeneration and Domain Biases: A Benchmark of LLMs on Advanced Mathematical Problem-Solving
di: Evstafev, Evgenii
Pubblicazione: (2025)
di: Evstafev, Evgenii
Pubblicazione: (2025)
The Paradox of Stochasticity: Limited Creativity and Computational Decoupling in Temperature-Varied LLM Outputs of Structured Fictional Data
di: Evstafev, Evgenii
Pubblicazione: (2025)
di: Evstafev, Evgenii
Pubblicazione: (2025)
Optimizing Humor Generation in Large Language Models: Temperature Configurations and Architectural Trade-offs
di: Evstafev, Evgenii
Pubblicazione: (2025)
di: Evstafev, Evgenii
Pubblicazione: (2025)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
di: Chua, James, et al.
Pubblicazione: (2025)
di: Chua, James, et al.
Pubblicazione: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
di: DeepSeek-AI, et al.
Pubblicazione: (2025)
di: DeepSeek-AI, et al.
Pubblicazione: (2025)
Brief analysis of DeepSeek R1 and its implications for Generative AI
di: Mercer, Sarah, et al.
Pubblicazione: (2025)
di: Mercer, Sarah, et al.
Pubblicazione: (2025)
Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
di: Spelda, Petr, et al.
Pubblicazione: (2025)
di: Spelda, Petr, et al.
Pubblicazione: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Memory Analysis on the Training Course of DeepSeek Models
di: Zhang, Ping, et al.
Pubblicazione: (2025)
di: Zhang, Ping, et al.
Pubblicazione: (2025)
A Review of DeepSeek Models' Key Innovative Techniques
di: Wang, Chengen, et al.
Pubblicazione: (2025)
di: Wang, Chengen, et al.
Pubblicazione: (2025)
Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
di: Zhao, Enbo, et al.
Pubblicazione: (2025)
di: Zhao, Enbo, et al.
Pubblicazione: (2025)
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
di: Menke, Antonio-Gabriel Chacón, et al.
Pubblicazione: (2025)
di: Menke, Antonio-Gabriel Chacón, et al.
Pubblicazione: (2025)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
di: Parmar, Manojkumar, et al.
Pubblicazione: (2025)
di: Parmar, Manojkumar, et al.
Pubblicazione: (2025)
Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis
di: Zhao, Kaikai, et al.
Pubblicazione: (2025)
di: Zhao, Kaikai, et al.
Pubblicazione: (2025)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
di: DeepSeek-AI, et al.
Pubblicazione: (2024)
di: DeepSeek-AI, et al.
Pubblicazione: (2024)
Benchmarking Multimodal Models for Fine-Grained Image Analysis: A Comparative Study Across Diverse Visual Features
di: Evstafev, Evgenii
Pubblicazione: (2025)
di: Evstafev, Evgenii
Pubblicazione: (2025)
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2025)
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2025)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
di: Yao, Jian, et al.
Pubblicazione: (2025)
di: Yao, Jian, et al.
Pubblicazione: (2025)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
di: Guo, Daya, et al.
Pubblicazione: (2024)
di: Guo, Daya, et al.
Pubblicazione: (2024)
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
di: DeepSeek-AI, et al.
Pubblicazione: (2024)
di: DeepSeek-AI, et al.
Pubblicazione: (2024)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
di: Yu, Zhouliang, et al.
Pubblicazione: (2025)
di: Yu, Zhouliang, et al.
Pubblicazione: (2025)
DeepSeek vs. ChatGPT vs. Claude: A Comparative Study for Scientific Computing and Scientific Machine Learning Tasks
di: Jiang, Qile, et al.
Pubblicazione: (2025)
di: Jiang, Qile, et al.
Pubblicazione: (2025)
DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey
di: Qiao, Yu, et al.
Pubblicazione: (2025)
di: Qiao, Yu, et al.
Pubblicazione: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
di: Wang, Ke, et al.
Pubblicazione: (2024)
di: Wang, Ke, et al.
Pubblicazione: (2024)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
di: Fan, Xiaoran, et al.
Pubblicazione: (2026)
di: Fan, Xiaoran, et al.
Pubblicazione: (2026)
Evaluating LLMs' Reasoning Over Ordered Procedural Steps
di: Anika, Adrita, et al.
Pubblicazione: (2025)
di: Anika, Adrita, et al.
Pubblicazione: (2025)
InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning
di: Zhang, Bo-Wen, et al.
Pubblicazione: (2024)
di: Zhang, Bo-Wen, et al.
Pubblicazione: (2024)
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
di: Xin, Huajian, et al.
Pubblicazione: (2024)
di: Xin, Huajian, et al.
Pubblicazione: (2024)
A Single Revision Step Improves Token-Efficient LLM Reasoning
di: Zhang, Yingchuan, et al.
Pubblicazione: (2026)
di: Zhang, Yingchuan, et al.
Pubblicazione: (2026)
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
di: Moell, Birger, et al.
Pubblicazione: (2025)
di: Moell, Birger, et al.
Pubblicazione: (2025)
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3
di: Sadik, Ahmed R., et al.
Pubblicazione: (2025)
di: Sadik, Ahmed R., et al.
Pubblicazione: (2025)
Mathematical Opportunities in Digital Twins (MATH-DT)
di: Antil, Harbir
Pubblicazione: (2024)
di: Antil, Harbir
Pubblicazione: (2024)
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2023)
di: Luccioni, Alexandra Sasha, et al.
Pubblicazione: (2023)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2025)
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2025)
Mixed-Precision Federated Learning via Multi-Precision Over-The-Air Aggregation
di: Yuan, Jinsheng, et al.
Pubblicazione: (2024)
di: Yuan, Jinsheng, et al.
Pubblicazione: (2024)
The Need for Speed: Pruning Transformers with One Recipe
di: Khaki, Samir, et al.
Pubblicazione: (2024)
di: Khaki, Samir, et al.
Pubblicazione: (2024)
Semantic Step Prediction: Multi-Step Latent Forecasting in LLM Reasoning Trajectories via Step Sampling
di: Yuan, Yidi
Pubblicazione: (2026)
di: Yuan, Yidi
Pubblicazione: (2026)
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
di: Sharma, Shubham, et al.
Pubblicazione: (2025)
di: Sharma, Shubham, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Token-by-Token Regeneration and Domain Biases: A Benchmark of LLMs on Advanced Mathematical Problem-Solving
di: Evstafev, Evgenii
Pubblicazione: (2025) -
The Paradox of Stochasticity: Limited Creativity and Computational Decoupling in Temperature-Varied LLM Outputs of Structured Fictional Data
di: Evstafev, Evgenii
Pubblicazione: (2025) -
Optimizing Humor Generation in Large Language Models: Temperature Configurations and Architectural Trade-offs
di: Evstafev, Evgenii
Pubblicazione: (2025) -
Are DeepSeek R1 And Other Reasoning Models More Faithful?
di: Chua, James, et al.
Pubblicazione: (2025) -
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
di: DeepSeek-AI, et al.
Pubblicazione: (2025)