R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Naseh, Ali, Chaudhari, Harsh, Roh, Jaechul, Wu, Mingshi, Oprea, Alina, Houmansadr, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025)
by: Suri, Anshuman, et al.
Published: (2025)
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
Identifying Models Behind Text-to-Image Leaderboards
by: Naseh, Ali, et al.
Published: (2026)
by: Naseh, Ali, et al.
Published: (2026)
Throttling Web Agents Using Reasoning Gates
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
Multilingual and Multi-Accent Jailbreaking of Audio LLMs
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
OverThink: Slowdown Attacks on Reasoning LLMs
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
OSLO: One-Shot Label-Only Membership Inference Attacks
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
CensorLess: Cost-Efficient Censorship Circumvention Through Serverless Cloud Functions
by: Kang, Dayeon, et al.
Published: (2026)
by: Kang, Dayeon, et al.
Published: (2026)
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
by: Parmar, Manojkumar, et al.
Published: (2025)
by: Parmar, Manojkumar, et al.
Published: (2025)
Diffence: Fencing Membership Privacy With Diffusion Models
by: Peng, Yuefeng, et al.
Published: (2023)
by: Peng, Yuefeng, et al.
Published: (2023)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
CensorLab: A Testbed for Censorship Experimentation
by: Sheffey, Jade, et al.
Published: (2024)
by: Sheffey, Jade, et al.
Published: (2024)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects
by: Rahman, Anichur, et al.
Published: (2025)
by: Rahman, Anichur, et al.
Published: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
Synthesizing Tight Privacy and Accuracy Bounds via Weighted Model Counting
by: Oakley, Lisa, et al.
Published: (2024)
by: Oakley, Lisa, et al.
Published: (2024)
Reconstruction of Personally Identifiable Information from Supervised Finetuned Models
by: Furukawa, Sae, et al.
Published: (2026)
by: Furukawa, Sae, et al.
Published: (2026)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
UTrace: Poisoning Forensics for Private Collaborative Learning
by: Rose, Evan, et al.
Published: (2024)
by: Rose, Evan, et al.
Published: (2024)
A Bayesian Approach to Membership Inference for Statistical Release
by: Oakley, Lisa, et al.
Published: (2026)
by: Oakley, Lisa, et al.
Published: (2026)
DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection
by: Ren, Junyu, et al.
Published: (2026)
by: Ren, Junyu, et al.
Published: (2026)
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024)
by: Chang, Yapei, et al.
Published: (2024)
User Inference Attacks on Large Language Models
by: Kandpal, Nikhil, et al.
Published: (2023)
by: Kandpal, Nikhil, et al.
Published: (2023)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
by: Wu, Xiaodong, et al.
Published: (2025)
by: Wu, Xiaodong, et al.
Published: (2025)
Membership Inference Attacks on Vision-Language-Action Models
by: Peng, Yuefeng, et al.
Published: (2026)
by: Peng, Yuefeng, et al.
Published: (2026)
FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models
by: Roh, Jaechul, et al.
Published: (2024)
by: Roh, Jaechul, et al.
Published: (2024)
Network-Level Prompt and Trait Leakage in Local Research Agents
by: Jeong, Hyejun, et al.
Published: (2025)
by: Jeong, Hyejun, et al.
Published: (2025)
RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation
by: Pham, Dzung, et al.
Published: (2023)
by: Pham, Dzung, et al.
Published: (2023)
Fake or Compromised? Making Sense of Malicious Clients in Federated Learning
by: Mozaffari, Hamid, et al.
Published: (2024)
by: Mozaffari, Hamid, et al.
Published: (2024)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
ProxyGPT: Enabling User Anonymity in LLM Chatbots via (Un)Trustworthy Volunteer Proxies
by: Pham, Dzung, et al.
Published: (2024)
by: Pham, Dzung, et al.
Published: (2024)
Toward a Principled Framework for Agent Safety Measurement
by: Lin, Shuyi, et al.
Published: (2026)
by: Lin, Shuyi, et al.
Published: (2026)
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
by: Xu, Zhiyuan, et al.
Published: (2025)
by: Xu, Zhiyuan, et al.
Published: (2025)
Similar Items
-
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
by: Naseh, Ali, et al.
Published: (2025) -
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
by: Naseh, Ali, et al.
Published: (2025) -
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025) -
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024) -
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)