Are DeepSeek R1 And Other Reasoning Models More Faithful?
Fuente:
arXiv
Saved in:
| Main Authors: | Chua, James, Evans, Owain |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
by: DeepSeek-AI, et al.
Published: (2025)
by: DeepSeek-AI, et al.
Published: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)
by: Chua, James, et al.
Published: (2026)
Brief analysis of DeepSeek R1 and its implications for Generative AI
by: Mercer, Sarah, et al.
Published: (2025)
by: Mercer, Sarah, et al.
Published: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Memory Analysis on the Training Course of DeepSeek Models
by: Zhang, Ping, et al.
Published: (2025)
by: Zhang, Ping, et al.
Published: (2025)
Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
by: Spelda, Petr, et al.
Published: (2025)
by: Spelda, Petr, et al.
Published: (2025)
A Review of DeepSeek Models' Key Innovative Techniques
by: Wang, Chengen, et al.
Published: (2025)
by: Wang, Chengen, et al.
Published: (2025)
Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH
by: Evstafev, Evgenii
Published: (2025)
by: Evstafev, Evgenii
Published: (2025)
Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
by: Zhao, Enbo, et al.
Published: (2025)
by: Zhao, Enbo, et al.
Published: (2025)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
by: Parmar, Manojkumar, et al.
Published: (2025)
by: Parmar, Manojkumar, et al.
Published: (2025)
Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis
by: Zhao, Kaikai, et al.
Published: (2025)
by: Zhao, Kaikai, et al.
Published: (2025)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
by: Guo, Daya, et al.
Published: (2024)
by: Guo, Daya, et al.
Published: (2024)
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Can Language Models Explain Their Own Classification Behavior?
by: Sherburn, Dane, et al.
Published: (2024)
by: Sherburn, Dane, et al.
Published: (2024)
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Are Reasoning Models More Prone to Hallucination?
by: Yao, Zijun, et al.
Published: (2025)
by: Yao, Zijun, et al.
Published: (2025)
Less Approximates More: Harmonizing Performance and Confidence Faithfulness via Hybrid Post-Training for High-Stakes Tasks
by: Ma, Haokai, et al.
Published: (2026)
by: Ma, Haokai, et al.
Published: (2026)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025)
by: Cloud, Alex, et al.
Published: (2025)
DeepSeek vs. ChatGPT vs. Claude: A Comparative Study for Scientific Computing and Scientific Machine Learning Tasks
by: Jiang, Qile, et al.
Published: (2025)
by: Jiang, Qile, et al.
Published: (2025)
DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey
by: Qiao, Yu, et al.
Published: (2025)
by: Qiao, Yu, et al.
Published: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
by: Xin, Huajian, et al.
Published: (2024)
by: Xin, Huajian, et al.
Published: (2024)
Mapping Faithful Reasoning in Language Models
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
by: Sharma, Shubham, et al.
Published: (2025)
by: Sharma, Shubham, et al.
Published: (2025)
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3
by: Sadik, Ahmed R., et al.
Published: (2025)
by: Sadik, Ahmed R., et al.
Published: (2025)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
A Comparison of DeepSeek and Other LLMs
by: Gao, Tianchen, et al.
Published: (2025)
by: Gao, Tianchen, et al.
Published: (2025)
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
by: Zhang, Chong, et al.
Published: (2025)
by: Zhang, Chong, et al.
Published: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
by: Jia, Jinghan, et al.
Published: (2026)
by: Jia, Jinghan, et al.
Published: (2026)
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
by: Guo, Yuhan, et al.
Published: (2025)
by: Guo, Yuhan, et al.
Published: (2025)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
by: Shao, Zhihong, et al.
Published: (2024)
by: Shao, Zhihong, et al.
Published: (2024)
Frugal, Flexible, Faithful: Causal Data Simulation via Frengression
by: Yang, Linying, et al.
Published: (2025)
by: Yang, Linying, et al.
Published: (2025)
Similar Items
-
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025) -
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
by: DeepSeek-AI, et al.
Published: (2025) -
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026) -
Brief analysis of DeepSeek R1 and its implications for Generative AI
by: Mercer, Sarah, et al.
Published: (2025) -
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)