Linearly Decoding Refused Knowledge in Aligned Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shrivastava, Aryan, Holtzman, Ari |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Know Thyself? On the Incapability and Implications of AI Self-Recognition
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2024)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2024)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
von: Cao, Lang
Veröffentlicht: (2023)
von: Cao, Lang
Veröffentlicht: (2023)
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
von: Li, Margaret, et al.
Veröffentlicht: (2024)
von: Li, Margaret, et al.
Veröffentlicht: (2024)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
AbsenceBench: Language Models Can't Tell What's Missing
von: Fu, Harvey Yiyun, et al.
Veröffentlicht: (2025)
von: Fu, Harvey Yiyun, et al.
Veröffentlicht: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
Measuring and Eliminating Refusals in Military Large Language Models
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
von: FitzGerald, Jack, et al.
Veröffentlicht: (2026)
Aligning Knowledge Graphs and Language Models for Factual Accuracy
von: Nishat, Nur A Zarin, et al.
Veröffentlicht: (2025)
von: Nishat, Nur A Zarin, et al.
Veröffentlicht: (2025)
Forking Paths in Neural Text Generation
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
von: Siu, Vincent, et al.
Veröffentlicht: (2025)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Approximately Aligned Decoding
von: Melcer, Daniel, et al.
Veröffentlicht: (2024)
von: Melcer, Daniel, et al.
Veröffentlicht: (2024)
MUSE: Machine Unlearning Six-Way Evaluation for Language Models
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
Advancing Reasoning in Large Language Models: Promising Methods and Approaches
von: Patil, Avinash, et al.
Veröffentlicht: (2025)
von: Patil, Avinash, et al.
Veröffentlicht: (2025)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2025)
Retrieval-Constrained Decoding Reveals Underestimated Parametric Knowledge in Language Models
von: Hamdani, Rajaa El, et al.
Veröffentlicht: (2025)
von: Hamdani, Rajaa El, et al.
Veröffentlicht: (2025)
LLM-Align: Utilizing Large Language Models for Entity Alignment in Knowledge Graphs
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
von: Mavi, John, et al.
Veröffentlicht: (2025)
von: Mavi, John, et al.
Veröffentlicht: (2025)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2025)
On Linearizing Structured Data in Encoder-Decoder Language Models: Insights from Text-to-SQL
von: Shao, Yutong, et al.
Veröffentlicht: (2024)
von: Shao, Yutong, et al.
Veröffentlicht: (2024)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
von: García-Ferrero, Iker, et al.
Veröffentlicht: (2025)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
Resource-Aware Arabic LLM Creation: Model Adaptation, Integration, and Multi-Domain Testing
von: Aryan, Prakash
Veröffentlicht: (2024)
von: Aryan, Prakash
Veröffentlicht: (2024)
Grammar-Aligned Decoding
von: Park, Kanghee, et al.
Veröffentlicht: (2024)
von: Park, Kanghee, et al.
Veröffentlicht: (2024)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
The Structure of Relation Decoding Linear Operators in Large Language Models
von: Christ, Miranda Anna, et al.
Veröffentlicht: (2025)
von: Christ, Miranda Anna, et al.
Veröffentlicht: (2025)
Distribution-Aligned Decoding for Efficient LLM Task Adaptation
von: Hu, Senkang, et al.
Veröffentlicht: (2025)
von: Hu, Senkang, et al.
Veröffentlicht: (2025)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
LLM-KT: Aligning Large Language Models with Knowledge Tracing using a Plug-and-Play Instruction
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
von: An, Zihao, et al.
Veröffentlicht: (2026)
von: An, Zihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Know Thyself? On the Incapability and Implications of AI Self-Recognition
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025) -
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
von: Tong, Zekai, et al.
Veröffentlicht: (2026) -
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2024) -
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
von: Cao, Lang
Veröffentlicht: (2023) -
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
von: Li, Margaret, et al.
Veröffentlicht: (2024)