The Violation State: Safety State Persistence in a Multimodal Language Model Interface
Fuente:
arXiv
Saved in:
| Main Author: | DeVilling, Bentley |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Polite Liar: Epistemic Pathology in Language Models
by: DeVilling, Bentley
Published: (2025)
by: DeVilling, Bentley
Published: (2025)
The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems
by: DeVilling, Bentley
Published: (2025)
by: DeVilling, Bentley
Published: (2025)
Assessing the State of AI Policy
by: DeFranco, Joanna F., et al.
Published: (2024)
by: DeFranco, Joanna F., et al.
Published: (2024)
Multimodal Modular Chain of Thoughts in Energy Performance Certificate Assessment
by: Peng, Zhen, et al.
Published: (2026)
by: Peng, Zhen, et al.
Published: (2026)
LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
by: Zhou, Guanghao, et al.
Published: (2026)
by: Zhou, Guanghao, et al.
Published: (2026)
Urban Safety Perception Through the Lens of Large Multimodal Models: A Persona-based Approach
by: Beneduce, Ciro, et al.
Published: (2025)
by: Beneduce, Ciro, et al.
Published: (2025)
Rethinking AI Literacy Education in Higher Education: Bridging Risk Perception and Responsible Adoption
by: Yu, Shasha, et al.
Published: (2026)
by: Yu, Shasha, et al.
Published: (2026)
Evaluating Psychological Safety of Large Language Models
by: Li, Xingxuan, et al.
Published: (2022)
by: Li, Xingxuan, et al.
Published: (2022)
Navigating the Edge with the State-of-the-Art Insights into Corner Case Identification and Generation for Enhanced Autonomous Vehicle Safety
by: Shimanuki, Gabriel Kenji Godoy, et al.
Published: (2025)
by: Shimanuki, Gabriel Kenji Godoy, et al.
Published: (2025)
AI Mimicry and Human Dignity: Chatbot Use as a Violation of Self-Respect
by: van der Rijt, Jan-Willem, et al.
Published: (2025)
by: van der Rijt, Jan-Willem, et al.
Published: (2025)
Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
by: Nakao, Mahiro, et al.
Published: (2026)
by: Nakao, Mahiro, et al.
Published: (2026)
The AI Pentad, the CHARME$^{2}$D Model, and an Assessment of Current-State AI Regulation
by: Gao, Di Kevin, et al.
Published: (2025)
by: Gao, Di Kevin, et al.
Published: (2025)
Defining and Evaluating Physical Safety for Large Language Models
by: Tang, Yung-Chen, et al.
Published: (2024)
by: Tang, Yung-Chen, et al.
Published: (2024)
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Case-based Reasoning Augmented Large Language Model Framework for Decision Making in Realistic Safety-Critical Driving Scenarios
by: Gan, Wenbin, et al.
Published: (2025)
by: Gan, Wenbin, et al.
Published: (2025)
Generative AI in Sociological Research: State of the Discipline
by: Alvero, AJ, et al.
Published: (2025)
by: Alvero, AJ, et al.
Published: (2025)
Can Large Language Models Simulate Human Responses? A Case Study of Stated Preference Experiments in the Context of Heating-related Choices
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education
by: Bewersdorff, Arne, et al.
Published: (2024)
by: Bewersdorff, Arne, et al.
Published: (2024)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
by: Scholefield, Rebecca, et al.
Published: (2025)
by: Scholefield, Rebecca, et al.
Published: (2025)
Phare: A Safety Probe for Large Language Models
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
by: Dobbe, Roel
Published: (2025)
by: Dobbe, Roel
Published: (2025)
Gauging Public Acceptance of Conditionally Automated Vehicles in the United States
by: Saravanos, Antonios, et al.
Published: (2024)
by: Saravanos, Antonios, et al.
Published: (2024)
From Text to Multimodality: Exploring the Evolution and Impact of Large Language Models in Medical Practice
by: Niu, Qian, et al.
Published: (2024)
by: Niu, Qian, et al.
Published: (2024)
Multimodal Safety Evaluation in Generative Agent Social Simulations
by: Vera, Alhim, et al.
Published: (2025)
by: Vera, Alhim, et al.
Published: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
Unveiling AI's Threats to Child Protection: Regulatory efforts to Criminalize AI-Generated CSAM and Emerging Children's Rights Violations
by: Kokolaki, Emmanouela, et al.
Published: (2025)
by: Kokolaki, Emmanouela, et al.
Published: (2025)
A Study on Individual Spatiotemporal Activity Generation Method Using MCP-Enhanced Chain-of-Thought Large Language Models
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
The Current State of AI Bias Bounties: An Overview of Existing Programmes and Research
by: Kucenko, Sergej, et al.
Published: (2025)
by: Kucenko, Sergej, et al.
Published: (2025)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Made-in China, Thinking in America:U.S. Values Persist in Chinese LLMs
by: Haslett, David, et al.
Published: (2025)
by: Haslett, David, et al.
Published: (2025)
Measuring the State of Open Science in Transportation Using Large Language Models
by: Ji, Junyi, et al.
Published: (2026)
by: Ji, Junyi, et al.
Published: (2026)
LLM Safety for Children
by: Rath, Prasanjit, et al.
Published: (2025)
by: Rath, Prasanjit, et al.
Published: (2025)
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
Mitigating Gambling-Like Risk-Taking Behaviors in Large Language Models: A Behavioral Economics Approach to AI Safety
by: Du, Y.
Published: (2025)
by: Du, Y.
Published: (2025)
Explainable Graph Neural Networks for Observation Impact Analysis in Atmospheric State Estimation
by: Jeon, Hyeon-Ju, et al.
Published: (2024)
by: Jeon, Hyeon-Ju, et al.
Published: (2024)
Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
On the Failure of Latent State Persistence in Large Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Similar Items
-
The Polite Liar: Epistemic Pathology in Language Models
by: DeVilling, Bentley
Published: (2025) -
The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems
by: DeVilling, Bentley
Published: (2025) -
Assessing the State of AI Policy
by: DeFranco, Joanna F., et al.
Published: (2024) -
Multimodal Modular Chain of Thoughts in Energy Performance Certificate Assessment
by: Peng, Zhen, et al.
Published: (2026) -
LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
by: Zhou, Guanghao, et al.
Published: (2026)