Refusal Behavior in Large Language Models: A Nonlinear Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hildebrandt, Fabian, Maier, Andreas, Krauss, Patrick, Schilling, Achim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Narrative Clustering in Large Language Models: A Layerwise Analysis of BERT
von: Banerjee, Awritrojit, et al.
Veröffentlicht: (2025)
von: Banerjee, Awritrojit, et al.
Veröffentlicht: (2025)
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
von: Krauss, Patrick, et al.
Veröffentlicht: (2024)
von: Krauss, Patrick, et al.
Veröffentlicht: (2024)
Analysis and Visualization of Linguistic Structures in Large Language Models: Neural Representations of Verb-Particle Constructions in BERT
von: Kissane, Hassane, et al.
Veröffentlicht: (2024)
von: Kissane, Hassane, et al.
Veröffentlicht: (2024)
Convergent Representations of Linguistic Constructions in Human and Artificial Neural Systems
von: Ramezani, Pegah, et al.
Veröffentlicht: (2026)
von: Ramezani, Pegah, et al.
Veröffentlicht: (2026)
Multi-Modal Cognitive Maps based on Neural Networks trained on Successor Representations
von: Stoewer, Paul, et al.
Veröffentlicht: (2023)
von: Stoewer, Paul, et al.
Veröffentlicht: (2023)
Analysis of Argument Structure Constructions in the Large Language Model BERT
von: Ramezani, Pegah, et al.
Veröffentlicht: (2024)
von: Ramezani, Pegah, et al.
Veröffentlicht: (2024)
Probing Internal Representations of Multi-Word Verbs in Large Language Models
von: Kissane, Hassane, et al.
Veröffentlicht: (2025)
von: Kissane, Hassane, et al.
Veröffentlicht: (2025)
Probing for Consciousness in Machines
von: Immertreu, Mathis, et al.
Veröffentlicht: (2024)
von: Immertreu, Mathis, et al.
Veröffentlicht: (2024)
Author-Specific Linguistic Patterns Unveiled: A Deep Learning Study on Word Class Distributions
von: Krauss, Patrick, et al.
Veröffentlicht: (2025)
von: Krauss, Patrick, et al.
Veröffentlicht: (2025)
Analysis of Argument Structure Constructions in a Deep Recurrent Language Model
von: Ramezani, Pegah, et al.
Veröffentlicht: (2024)
von: Ramezani, Pegah, et al.
Veröffentlicht: (2024)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
Measuring and Eliminating Refusals in Military Large Language Models
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
von: Cao, Lang
Veröffentlicht: (2023)
von: Cao, Lang
Veröffentlicht: (2023)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
Classifying German Language Proficiency Levels Using Large Language Models
von: Ahlers, Elias-Leander, et al.
Veröffentlicht: (2025)
von: Ahlers, Elias-Leander, et al.
Veröffentlicht: (2025)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
von: von Recum, Alexander, et al.
Veröffentlicht: (2024)
von: von Recum, Alexander, et al.
Veröffentlicht: (2024)
A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
von: Linder, Noa, et al.
Veröffentlicht: (2026)
von: Linder, Noa, et al.
Veröffentlicht: (2026)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Nonlinear Neural Dynamics and Classification Accuracy in Reservoir Computing
von: Metzner, Claus, et al.
Veröffentlicht: (2024)
von: Metzner, Claus, et al.
Veröffentlicht: (2024)
Scan-do Attitude: Towards Autonomous CT Protocol Management using a Large Language Model Agent
von: Kang, Xingjian, et al.
Veröffentlicht: (2025)
von: Kang, Xingjian, et al.
Veröffentlicht: (2025)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective
von: Mao, Liyuan, et al.
Veröffentlicht: (2026)
von: Mao, Liyuan, et al.
Veröffentlicht: (2026)
Behavioral Fingerprinting of Large Language Models
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
The Mercurial Top-Level Ontology of Large Language Models
von: Köhler, Nele, et al.
Veröffentlicht: (2024)
von: Köhler, Nele, et al.
Veröffentlicht: (2024)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
Reasoning in Large Language Models: A Geometric Perspective
von: Cosentino, Romain, et al.
Veröffentlicht: (2024)
von: Cosentino, Romain, et al.
Veröffentlicht: (2024)
Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language
von: Immertreu, Mathis, et al.
Veröffentlicht: (2026)
von: Immertreu, Mathis, et al.
Veröffentlicht: (2026)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
A Multi-Perspective Analysis of Memorization in Large Language Models
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
von: Chen, Bowen, et al.
Veröffentlicht: (2024)
Visuospatial Perspective Taking in Multimodal Language Models
von: Prunty, Jonathan, et al.
Veröffentlicht: (2026)
von: Prunty, Jonathan, et al.
Veröffentlicht: (2026)
Investigating the Robustness of Deductive Reasoning with Large Language Models
von: Hoppe, Fabian, et al.
Veröffentlicht: (2025)
von: Hoppe, Fabian, et al.
Veröffentlicht: (2025)
Where Do Reasoning Models Refuse?
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
The System Hallucination Scale (SHS): A Minimal yet Effective Human-Centered Instrument for Evaluating Hallucination-Related Behavior in Large Language Models
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Exploring Narrative Clustering in Large Language Models: A Layerwise Analysis of BERT
von: Banerjee, Awritrojit, et al.
Veröffentlicht: (2025) -
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
von: Krauss, Patrick, et al.
Veröffentlicht: (2024) -
Analysis and Visualization of Linguistic Structures in Large Language Models: Neural Representations of Verb-Particle Constructions in BERT
von: Kissane, Hassane, et al.
Veröffentlicht: (2024) -
Convergent Representations of Linguistic Constructions in Human and Artificial Neural Systems
von: Ramezani, Pegah, et al.
Veröffentlicht: (2026) -
Multi-Modal Cognitive Maps based on Neural Networks trained on Successor Representations
von: Stoewer, Paul, et al.
Veröffentlicht: (2023)