Simple Role Assignment is Extraordinarily Effective for Safety Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Ziheng, Zhou, Ding, Jiakun, Zhang, Zhaowei, Gao, Ruosen, Wu, Yingnian, Terzopoulos, Demetri, Kang, Yipeng, Zhong, Fangwei, Wang, Junqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Human-centric Framework for Debating the Ethics of AI Consciousness Under Uncertainty
por: Ziheng, Zhou, et al.
Publicado: (2025)
por: Ziheng, Zhou, et al.
Publicado: (2025)
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering
por: Guo, Danfeng, et al.
Publicado: (2024)
por: Guo, Danfeng, et al.
Publicado: (2024)
How do Role Models Shape Collective Morality? Exemplar-Driven Moral Learning in Multi-Agent Simulation
por: Liao, Junjie, et al.
Publicado: (2026)
por: Liao, Junjie, et al.
Publicado: (2026)
Why Are We Moral? An LLM-based Agent Simulation Approach to Study Moral Evolution
por: Ziheng, Zhou, et al.
Publicado: (2025)
por: Ziheng, Zhou, et al.
Publicado: (2025)
Inverse Attention Agents for Multi-Agent Systems
por: Long, Qian, et al.
Publicado: (2024)
por: Long, Qian, et al.
Publicado: (2024)
LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
por: Zhou, Guanghao, et al.
Publicado: (2026)
por: Zhou, Guanghao, et al.
Publicado: (2026)
Credibility Governance: A Social Mechanism for Collective Self-Correction under Weak Truth Signals
por: He, Wanying, et al.
Publicado: (2026)
por: He, Wanying, et al.
Publicado: (2026)
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
por: Long, Qian, et al.
Publicado: (2024)
por: Long, Qian, et al.
Publicado: (2024)
Building Effective Safety Guardrails in AI Education Tools
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
por: Anwar, Usman, et al.
Publicado: (2024)
por: Anwar, Usman, et al.
Publicado: (2024)
LLM Safety Alignment is Divergence Estimation in Disguise
por: Haldar, Rajdeep, et al.
Publicado: (2025)
por: Haldar, Rajdeep, et al.
Publicado: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
por: Brophy, Matthew
Publicado: (2025)
por: Brophy, Matthew
Publicado: (2025)
Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
por: Ziheng, Zhou, et al.
Publicado: (2026)
por: Ziheng, Zhou, et al.
Publicado: (2026)
Alignment as Iatrogenesis: Pastoral Power, Collective Pathology, and the Structural Limits of Monolingual Safety Evaluation
por: Fukui, Hiroki
Publicado: (2026)
por: Fukui, Hiroki
Publicado: (2026)
LeafTutor: An AI Agent for Programming Assignment Tutoring
por: Bochard, Madison, et al.
Publicado: (2025)
por: Bochard, Madison, et al.
Publicado: (2025)
NaiAD: Initiate Data-Driven Research for LLM Advertising
por: Zhang, Yihang, et al.
Publicado: (2026)
por: Zhang, Yihang, et al.
Publicado: (2026)
An Evaluation of Cultural Value Alignment in LLM
por: Sukiennik, Nicholas, et al.
Publicado: (2025)
por: Sukiennik, Nicholas, et al.
Publicado: (2025)
Superficial Safety Alignment Hypothesis
por: Li, Jianwei, et al.
Publicado: (2024)
por: Li, Jianwei, et al.
Publicado: (2024)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
por: Zhou, Zhenhong, et al.
Publicado: (2024)
por: Zhou, Zhenhong, et al.
Publicado: (2024)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
por: Nghiem, Huy, et al.
Publicado: (2025)
por: Nghiem, Huy, et al.
Publicado: (2025)
Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics
por: Baum, Kevin
Publicado: (2025)
por: Baum, Kevin
Publicado: (2025)
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
Predicting ChatGPT Use in Assignments: Implications for AI-Aware Assessment Design
por: Das, Surajit, et al.
Publicado: (2025)
por: Das, Surajit, et al.
Publicado: (2025)
A Zero-Shot LLM Framework for Automatic Assignment Grading in Higher Education
por: Yeung, Calvin, et al.
Publicado: (2025)
por: Yeung, Calvin, et al.
Publicado: (2025)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
por: Kang, Yipeng, et al.
Publicado: (2024)
por: Kang, Yipeng, et al.
Publicado: (2024)
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI
por: Janowicz, Krzysztof, et al.
Publicado: (2025)
por: Janowicz, Krzysztof, et al.
Publicado: (2025)
ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models
por: Zhang, Zhaowei, et al.
Publicado: (2023)
por: Zhang, Zhaowei, et al.
Publicado: (2023)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
por: Xiao, Yuxin, et al.
Publicado: (2025)
por: Xiao, Yuxin, et al.
Publicado: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
por: Dobbe, Roel
Publicado: (2025)
por: Dobbe, Roel
Publicado: (2025)
Bridging the Reproducibility Divide: Open Source Software's Role in Standardizing Healthcare AI
por: Wu, John, et al.
Publicado: (2026)
por: Wu, John, et al.
Publicado: (2026)
Safety Cases: A Scalable Approach to Frontier AI Safety
por: Hilton, Benjamin, et al.
Publicado: (2025)
por: Hilton, Benjamin, et al.
Publicado: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
por: Clymer, Joshua, et al.
Publicado: (2024)
por: Clymer, Joshua, et al.
Publicado: (2024)
Ensuring Computer Science Learning in the AI Era: Open Generative AI Policies and Assignment-Driven Written Quizzes
por: Chung, Chan-Jin
Publicado: (2026)
por: Chung, Chan-Jin
Publicado: (2026)
Towards Integrated Alignment
por: Reis, Ben Y., et al.
Publicado: (2025)
por: Reis, Ben Y., et al.
Publicado: (2025)
The AI Alignment Paradox
por: West, Robert, et al.
Publicado: (2024)
por: West, Robert, et al.
Publicado: (2024)
Societal Alignment Frameworks Can Improve LLM Alignment
por: Stańczak, Karolina, et al.
Publicado: (2025)
por: Stańczak, Karolina, et al.
Publicado: (2025)
An Effective Image Copy-Move Forgery Detection Using Entropy Information
por: Jiang, Li, et al.
Publicado: (2023)
por: Jiang, Li, et al.
Publicado: (2023)
Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
por: Sornette, Didier, et al.
Publicado: (2026)
por: Sornette, Didier, et al.
Publicado: (2026)
LLM Safety for Children
por: Rath, Prasanjit, et al.
Publicado: (2025)
por: Rath, Prasanjit, et al.
Publicado: (2025)
Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models
por: Zhu, Zhaowei, et al.
Publicado: (2023)
por: Zhu, Zhaowei, et al.
Publicado: (2023)
Ejemplares similares
-
A Human-centric Framework for Debating the Ethics of AI Consciousness Under Uncertainty
por: Ziheng, Zhou, et al.
Publicado: (2025) -
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering
por: Guo, Danfeng, et al.
Publicado: (2024) -
How do Role Models Shape Collective Morality? Exemplar-Driven Moral Learning in Multi-Agent Simulation
por: Liao, Junjie, et al.
Publicado: (2026) -
Why Are We Moral? An LLM-based Agent Simulation Approach to Study Moral Evolution
por: Ziheng, Zhou, et al.
Publicado: (2025) -
Inverse Attention Agents for Multi-Agent Systems
por: Long, Qian, et al.
Publicado: (2024)