Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xuying, Li, Zhuo, Kosuga, Yuji, Bian, Victor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Precision Knowledge Editing: Enhancing Safety in Large Language Models
von: Li, Xuying, et al.
Veröffentlicht: (2024)
von: Li, Xuying, et al.
Veröffentlicht: (2024)
Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach
von: Li, Xuying, et al.
Veröffentlicht: (2025)
von: Li, Xuying, et al.
Veröffentlicht: (2025)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
von: Shu, Huizhen, et al.
Veröffentlicht: (2025)
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation
von: Li, Xuying, et al.
Veröffentlicht: (2024)
von: Li, Xuying, et al.
Veröffentlicht: (2024)
Safety Evaluation of DeepSeek Models in Chinese Contexts
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
DeepSeek-V3 Technical Report
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
Mixture of Tunable Experts -- Behavior Modification of DeepSeek-R1 at Inference Time
von: Dahlke, Robert, et al.
Veröffentlicht: (2025)
von: Dahlke, Robert, et al.
Veröffentlicht: (2025)
A Comparison of DeepSeek and Other LLMs
von: Gao, Tianchen, et al.
Veröffentlicht: (2025)
von: Gao, Tianchen, et al.
Veröffentlicht: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
Explainable Sentiment Analysis with DeepSeek-R1: Performance, Efficiency, and Few-Shot Learning
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
von: Parmar, Manojkumar, et al.
Veröffentlicht: (2025)
von: Parmar, Manojkumar, et al.
Veröffentlicht: (2025)
An evaluation of DeepSeek Models in Biomedical Natural Language Processing
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
Reasoning and the Trusting Behavior of DeepSeek and GPT: An Experiment Revealing Hidden Fault Lines in Large Language Models
von: Li, Rubing, et al.
Veröffentlicht: (2025)
von: Li, Rubing, et al.
Veröffentlicht: (2025)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
LLMs in Disease Diagnosis: A Comparative Study of DeepSeek-R1 and O3 Mini Across Chronic Health Conditions
von: Gupta, Gaurav Kumar, et al.
Veröffentlicht: (2025)
von: Gupta, Gaurav Kumar, et al.
Veröffentlicht: (2025)
DeepSeek performs better than other Large Language Models in Dental Cases
von: Zhang, Hexian, et al.
Veröffentlicht: (2025)
von: Zhang, Hexian, et al.
Veröffentlicht: (2025)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking
von: Li, Baixuan, et al.
Veröffentlicht: (2025)
von: Li, Baixuan, et al.
Veröffentlicht: (2025)
Emotion-Aware Embedding Fusion in LLMs (Flan-T5, LLAMA 2, DeepSeek-R1, and ChatGPT 4) for Intelligent Response Generation
von: Rasool, Abdur, et al.
Veröffentlicht: (2024)
von: Rasool, Abdur, et al.
Veröffentlicht: (2024)
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
von: Ji, Tao, et al.
Veröffentlicht: (2025)
von: Ji, Tao, et al.
Veröffentlicht: (2025)
DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models
von: Ye, Jiancheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiancheng, et al.
Veröffentlicht: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
Comparative Analysis of OpenAI GPT-4o and DeepSeek R1 for Scientific Text Categorization Using Prompt Engineering
von: Maiti, Aniruddha, et al.
Veröffentlicht: (2025)
von: Maiti, Aniruddha, et al.
Veröffentlicht: (2025)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
von: Sands, Brendan, et al.
Veröffentlicht: (2025)
von: Sands, Brendan, et al.
Veröffentlicht: (2025)
Comparative Evaluation of ChatGPT and DeepSeek Across Key NLP Tasks: Strengths, Weaknesses, and Domain-Specific Performance
von: Etaiwi, Wael, et al.
Veröffentlicht: (2025)
von: Etaiwi, Wael, et al.
Veröffentlicht: (2025)
Evaluating the Performance of AI Text Detectors, Few-Shot and Chain-of-Thought Prompting Using DeepSeek Generated Text
von: Alshammari, Hulayyil, et al.
Veröffentlicht: (2025)
von: Alshammari, Hulayyil, et al.
Veröffentlicht: (2025)
A comprehensive study of LLM-based argument classification: from Llama through DeepSeek to GPT-5.2
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
FloorPlan-DeepSeek (FPDS): A multimodal approach to floorplan generation using vector-based next room prediction
von: Yin, Jun, et al.
Veröffentlicht: (2025)
von: Yin, Jun, et al.
Veröffentlicht: (2025)
Bridging Technology and Humanities: Evaluating the Impact of Large Language Models on Social Sciences Research with DeepSeek-R1
von: Gu, Peiran, et al.
Veröffentlicht: (2025)
von: Gu, Peiran, et al.
Veröffentlicht: (2025)
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
von: Xin, Huajian, et al.
Veröffentlicht: (2024)
von: Xin, Huajian, et al.
Veröffentlicht: (2024)
AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
von: Rui, Shaohao, et al.
Veröffentlicht: (2025)
von: Rui, Shaohao, et al.
Veröffentlicht: (2025)
DeepSeek reshaping healthcare in China's tertiary hospitals
von: Chen, Jishizhan, et al.
Veröffentlicht: (2025)
von: Chen, Jishizhan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Precision Knowledge Editing: Enhancing Safety in Large Language Models
von: Li, Xuying, et al.
Veröffentlicht: (2024) -
Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach
von: Li, Xuying, et al.
Veröffentlicht: (2025) -
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
von: Shu, Huizhen, et al.
Veröffentlicht: (2025) -
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
von: Zhang, Yichi, et al.
Veröffentlicht: (2025) -
Targeting the Core: A Simple and Effective Method to Attack RAG-based Agents via Direct LLM Manipulation
von: Li, Xuying, et al.
Veröffentlicht: (2024)