The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wollschläger, Tom, Elstner, Jannes, Geisler, Simon, Cohen-Addad, Vincent, Günnemann, Stephan, Gasteiger, Johannes |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
von: Geisler, Simon, et al.
Veröffentlicht: (2025)
von: Geisler, Simon, et al.
Veröffentlicht: (2025)
Attacking Large Language Models with Projected Gradient Descent
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
Consistency Training while Mitigating Obfuscation via Rate Matching
von: Imran, Sohaib, et al.
Veröffentlicht: (2026)
von: Imran, Sohaib, et al.
Veröffentlicht: (2026)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Solving an Open Problem in Theoretical Physics using AI-Assisted Discovery
von: Brenner, Michael P., et al.
Veröffentlicht: (2026)
von: Brenner, Michael P., et al.
Veröffentlicht: (2026)
Measuring and Eliminating Refusals in Military Large Language Models
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
Algorithmic Thinking Theory
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2025)
von: Bateni, MohammadHossein, et al.
Veröffentlicht: (2025)
Long-Range Graph Wavelet Networks
von: Guerranti, Filippo, et al.
Veröffentlicht: (2025)
von: Guerranti, Filippo, et al.
Veröffentlicht: (2025)
Spatio-Spectral Graph Neural Networks
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
Measuring Representation Robustness in Large Language Models for Geometry
von: Jawandhia, Vedant, et al.
Veröffentlicht: (2026)
von: Jawandhia, Vedant, et al.
Veröffentlicht: (2026)
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
von: Park, Kiho, et al.
Veröffentlicht: (2024)
von: Park, Kiho, et al.
Veröffentlicht: (2024)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
von: Cao, Lang
Veröffentlicht: (2023)
von: Cao, Lang
Veröffentlicht: (2023)
Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains
von: Fonseca, Marcio, et al.
Veröffentlicht: (2023)
von: Fonseca, Marcio, et al.
Veröffentlicht: (2023)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
The Linear Representation Hypothesis and the Geometry of Large Language Models
von: Park, Kiho, et al.
Veröffentlicht: (2023)
von: Park, Kiho, et al.
Veröffentlicht: (2023)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
Emotion Concepts and their Function in a Large Language Model
von: Sofroniew, Nicholas, et al.
Veröffentlicht: (2026)
von: Sofroniew, Nicholas, et al.
Veröffentlicht: (2026)
Evaluating and Understanding Scheming Propensity in LLM Agents
von: Hopman, Mia, et al.
Veröffentlicht: (2026)
von: Hopman, Mia, et al.
Veröffentlicht: (2026)
A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
von: Linder, Noa, et al.
Veröffentlicht: (2026)
von: Linder, Noa, et al.
Veröffentlicht: (2026)
From Knowledge to Treatment: Large Language Model Assisted Biomedical Concept Representation for Drug Repurposing
von: Xiang, Chengrui, et al.
Veröffentlicht: (2025)
von: Xiang, Chengrui, et al.
Veröffentlicht: (2025)
Context Structure Reshapes the Representational Geometry of Language Models
von: Hosseini, Eghbal A., et al.
Veröffentlicht: (2026)
von: Hosseini, Eghbal A., et al.
Veröffentlicht: (2026)
Unlocking Reasoning Capability on Machine Translation in Large Language Models
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
Large Language Model for Patent Concept Generation
von: Ren, Runtao, et al.
Veröffentlicht: (2024)
von: Ren, Runtao, et al.
Veröffentlicht: (2024)
Explainable Graph Neural Networks Under Fire
von: Li, Zhong, et al.
Veröffentlicht: (2024)
von: Li, Zhong, et al.
Veröffentlicht: (2024)
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
von: Scholten, Yan, et al.
Veröffentlicht: (2024)
von: Scholten, Yan, et al.
Veröffentlicht: (2024)
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
von: Borah, Abhilekh, et al.
Veröffentlicht: (2025)
von: Borah, Abhilekh, et al.
Veröffentlicht: (2025)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
Language-Independent Representations Improve Zero-Shot Summarization
von: Solovyev, Vladimir, et al.
Veröffentlicht: (2024)
von: Solovyev, Vladimir, et al.
Veröffentlicht: (2024)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
Identifying Linear Relational Concepts in Large Language Models
von: Chanin, David, et al.
Veröffentlicht: (2023)
von: Chanin, David, et al.
Veröffentlicht: (2023)
ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
von: Lermen, Simon, et al.
Veröffentlicht: (2024)
von: Lermen, Simon, et al.
Veröffentlicht: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Language Independent Stance Detection: Social Interaction-based Embeddings and Large Language Models
von: de Landa, Joseba Fernandez, et al.
Veröffentlicht: (2022)
von: de Landa, Joseba Fernandez, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
von: Geisler, Simon, et al.
Veröffentlicht: (2025) -
Attacking Large Language Models with Projected Gradient Descent
von: Geisler, Simon, et al.
Veröffentlicht: (2024) -
Consistency Training while Mitigating Obfuscation via Rate Matching
von: Imran, Sohaib, et al.
Veröffentlicht: (2026) -
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025) -
Solving an Open Problem in Theoretical Physics using AI-Assisted Discovery
von: Brenner, Michael P., et al.
Veröffentlicht: (2026)