How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Menke, Antonio-Gabriel Chacón, Tan, Phan Xuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917982330945536
author Menke, Antonio-Gabriel Chacón
Tan, Phan Xuan
author_facet Menke, Antonio-Gabriel Chacón
Tan, Phan Xuan
contents Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-critique mechanism on small, uncensored 7-9B parameter models: DeepSeek-R1-8B, Gemma-2-9B, Llama 3.1-8B, and Qwen2.5-7B. We show that while Llama-based models exhibited significant harm reduction through self-critique, other architectures demonstrated less improvement in harm detection after abliteration. These results suggest CAI's effectiveness may vary depending on model architecture and reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
Menke, Antonio-Gabriel Chacón
Tan, Phan Xuan
Machine Learning
Artificial Intelligence
Computers and Society
Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-critique mechanism on small, uncensored 7-9B parameter models: DeepSeek-R1-8B, Gemma-2-9B, Llama 3.1-8B, and Qwen2.5-7B. We show that while Llama-based models exhibited significant harm reduction through self-critique, other architectures demonstrated less improvement in harm detection after abliteration. These results suggest CAI's effectiveness may vary depending on model architecture and reasoning capabilities.
title How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers
topic Machine Learning
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2503.17365