Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Spelda, Petr, Stritecky, Vit |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Security practices in AI development
par: Spelda, Petr, et autres
Publié: (2025)
par: Spelda, Petr, et autres
Publié: (2025)
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
par: Menke, Antonio-Gabriel Chacón, et autres
Publié: (2025)
par: Menke, Antonio-Gabriel Chacón, et autres
Publié: (2025)
Brief analysis of DeepSeek R1 and its implications for Generative AI
par: Mercer, Sarah, et autres
Publié: (2025)
par: Mercer, Sarah, et autres
Publié: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
par: Naseh, Ali, et autres
Publié: (2025)
par: Naseh, Ali, et autres
Publié: (2025)
Institutional Trust and the Domestic AI Advantage: Evidence from DeepSeek and ChatGPT Users in China
par: Huang, Jiashen, et autres
Publié: (2026)
par: Huang, Jiashen, et autres
Publié: (2026)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
par: DeepSeek-AI, et autres
Publié: (2025)
par: DeepSeek-AI, et autres
Publié: (2025)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
par: Parmar, Manojkumar, et autres
Publié: (2025)
par: Parmar, Manojkumar, et autres
Publié: (2025)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
par: Chua, James, et autres
Publié: (2025)
par: Chua, James, et autres
Publié: (2025)
Is Power-Seeking AI an Existential Risk?
par: Carlsmith, Joseph
Publié: (2022)
par: Carlsmith, Joseph
Publié: (2022)
DeepSeek reshaping healthcare in China's tertiary hospitals
par: Chen, Jishizhan, et autres
Publié: (2025)
par: Chen, Jishizhan, et autres
Publié: (2025)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
par: Wang, Qian, et autres
Publié: (2025)
par: Wang, Qian, et autres
Publié: (2025)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
par: DeepSeek-AI, et autres
Publié: (2024)
par: DeepSeek-AI, et autres
Publié: (2024)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
par: Zhang, Wenjing, et autres
Publié: (2025)
par: Zhang, Wenjing, et autres
Publié: (2025)
Australian Bushfire Intelligence with AI-Driven Environmental Analytics
par: Jois, Tanvi, et autres
Publié: (2026)
par: Jois, Tanvi, et autres
Publié: (2026)
Bridging Technology and Humanities: Evaluating the Impact of Large Language Models on Social Sciences Research with DeepSeek-R1
par: Gu, Peiran, et autres
Publié: (2025)
par: Gu, Peiran, et autres
Publié: (2025)
Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis
par: Zhao, Kaikai, et autres
Publié: (2025)
par: Zhao, Kaikai, et autres
Publié: (2025)
From Transparency to Accountability and Back: A Discussion of Access and Evidence in AI Auditing
par: Cen, Sarah H., et autres
Publié: (2024)
par: Cen, Sarah H., et autres
Publié: (2024)
Memory Analysis on the Training Course of DeepSeek Models
par: Zhang, Ping, et autres
Publié: (2025)
par: Zhang, Ping, et autres
Publié: (2025)
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
par: Islam, Chashi Mahiul, et autres
Publié: (2025)
par: Islam, Chashi Mahiul, et autres
Publié: (2025)
Semantic Risk Scoring of Aggregated Metrics: An AI-Driven Approach for Healthcare Data Governance
par: Ahmed, Mohammed Omer Shakeel
Publié: (2026)
par: Ahmed, Mohammed Omer Shakeel
Publié: (2026)
AI-Driven Strategies for Reducing Student Withdrawal -- A Study of EMU Student Stopout
par: Zhao, Yan, et autres
Publié: (2024)
par: Zhao, Yan, et autres
Publié: (2024)
Can LLMs Assist Computer Education? an Empirical Case Study of DeepSeek
par: Xiao, Dongfu, et autres
Publié: (2025)
par: Xiao, Dongfu, et autres
Publié: (2025)
A Review of DeepSeek Models' Key Innovative Techniques
par: Wang, Chengen, et autres
Publié: (2025)
par: Wang, Chengen, et autres
Publié: (2025)
Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma
par: Aydin, Omer, et autres
Publié: (2025)
par: Aydin, Omer, et autres
Publié: (2025)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
par: Guo, Daya, et autres
Publié: (2024)
par: Guo, Daya, et autres
Publié: (2024)
SweetDeep: A Wearable AI Solution for Real-Time Non-Invasive Diabetes Screening
par: Henriques, Ian, et autres
Publié: (2025)
par: Henriques, Ian, et autres
Publié: (2025)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
par: Qiu, Peiran, et autres
Publié: (2025)
par: Qiu, Peiran, et autres
Publié: (2025)
Can AI be a Teaching Partner? Evaluating ChatGPT, Gemini, and DeepSeek across Three Teaching Strategies
par: de Souza, Talita de Paula Cypriano, et autres
Publié: (2026)
par: de Souza, Talita de Paula Cypriano, et autres
Publié: (2026)
Examining the Relationship between Scientific Publishing Activity and Hype-Driven Financial Bubbles: A Comparison of the Dot-Com and AI Eras
par: Chelikavada, Aksheytha, et autres
Publié: (2025)
par: Chelikavada, Aksheytha, et autres
Publié: (2025)
AI Oversight and Human Mistakes: Evidence from Centre Court
par: Almog, David, et autres
Publié: (2024)
par: Almog, David, et autres
Publié: (2024)
Transparent and Fair Profiling in Employment Services: Evidence from Switzerland
par: Räz, Tim
Publié: (2025)
par: Räz, Tim
Publié: (2025)
Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
par: Zhao, Enbo, et autres
Publié: (2025)
par: Zhao, Enbo, et autres
Publié: (2025)
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
par: Rabanser, Stephan
Publié: (2025)
par: Rabanser, Stephan
Publié: (2025)
Benchmarking Stochastic Approximation Algorithms for Fairness-Constrained Training of Deep Neural Networks
par: Kliachkin, Andrii, et autres
Publié: (2025)
par: Kliachkin, Andrii, et autres
Publié: (2025)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
par: Gringras, David
Publié: (2026)
par: Gringras, David
Publié: (2026)
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
par: Jurenka, Irina, et autres
Publié: (2024)
par: Jurenka, Irina, et autres
Publié: (2024)
Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH
par: Evstafev, Evgenii
Publié: (2025)
par: Evstafev, Evgenii
Publié: (2025)
Emotion-Aware Embedding Fusion in LLMs (Flan-T5, LLAMA 2, DeepSeek-R1, and ChatGPT 4) for Intelligent Response Generation
par: Rasool, Abdur, et autres
Publié: (2024)
par: Rasool, Abdur, et autres
Publié: (2024)
From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution
par: Koch, Bernard J., et autres
Publié: (2024)
par: Koch, Bernard J., et autres
Publié: (2024)
Extinction Risks from AI: Invisible to Science?
par: Kovarik, Vojtech, et autres
Publié: (2024)
par: Kovarik, Vojtech, et autres
Publié: (2024)
Documents similaires
-
Security practices in AI development
par: Spelda, Petr, et autres
Publié: (2025) -
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
par: Menke, Antonio-Gabriel Chacón, et autres
Publié: (2025) -
Brief analysis of DeepSeek R1 and its implications for Generative AI
par: Mercer, Sarah, et autres
Publié: (2025) -
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
par: Naseh, Ali, et autres
Publié: (2025) -
Institutional Trust and the Domestic AI Advantage: Evidence from DeepSeek and ChatGPT Users in China
par: Huang, Jiashen, et autres
Publié: (2026)