Saved in:
| Main Authors: | Yamaguchi, Kureha, Etheridge, Benjamin, Arditi, Andy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.03167 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Refusal in Language Models Is Mediated by a Single Direction
by: Arditi, Andy, et al.
Published: (2024)
by: Arditi, Andy, et al.
Published: (2024)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
by: Obeso, Oscar, et al.
Published: (2025)
by: Obeso, Oscar, et al.
Published: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Programming Refusal with Conditional Activation Steering
by: Lee, Bruce W., et al.
Published: (2024)
by: Lee, Bruce W., et al.
Published: (2024)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
by: Wollschläger, Tom, et al.
Published: (2025)
by: Wollschläger, Tom, et al.
Published: (2025)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
by: Turpin, Miles, et al.
Published: (2025)
by: Turpin, Miles, et al.
Published: (2025)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
by: Kang, Deokhyung, et al.
Published: (2025)
by: Kang, Deokhyung, et al.
Published: (2025)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
by: Abrams, Mitchell, et al.
Published: (2026)
by: Abrams, Mitchell, et al.
Published: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
by: Hu, Xulin, et al.
Published: (2026)
by: Hu, Xulin, et al.
Published: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
by: Jiang, Eric Hanchen, et al.
Published: (2025)
by: Jiang, Eric Hanchen, et al.
Published: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
by: Frank, Gregory N.
Published: (2026)
by: Frank, Gregory N.
Published: (2026)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
by: Cheng, Stephen, et al.
Published: (2026)
by: Cheng, Stephen, et al.
Published: (2026)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
Do LLMs Encode Functional Importance of Reasoning Tokens?
by: Singh, Janvijay, et al.
Published: (2026)
by: Singh, Janvijay, et al.
Published: (2026)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
by: Du, Hongzhe, et al.
Published: (2025)
by: Du, Hongzhe, et al.
Published: (2025)
Steering Large Reasoning Models towards Concise Reasoning via Flow Matching
by: Li, Yawei, et al.
Published: (2026)
by: Li, Yawei, et al.
Published: (2026)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026)
by: Beniwal, Himanshu, et al.
Published: (2026)
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
by: Cui, Yingqian, et al.
Published: (2026)
by: Cui, Yingqian, et al.
Published: (2026)
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024)
by: Zellinger, Lena, et al.
Published: (2024)
Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
by: Li, Chengshu, et al.
Published: (2023)
by: Li, Chengshu, et al.
Published: (2023)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
by: Gopalakrishnan, Anand, et al.
Published: (2025)
by: Gopalakrishnan, Anand, et al.
Published: (2025)
Where does output diversity collapse in post-training?
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
by: Xia, Wei, et al.
Published: (2026)
by: Xia, Wei, et al.
Published: (2026)
Do Androids Know They're Only Dreaming of Electric Sheep?
by: CH-Wang, Sky, et al.
Published: (2023)
by: CH-Wang, Sky, et al.
Published: (2023)
Mini-Giants: "Small" Language Models and Open Source Win-Win
by: Zhou, Zhengping, et al.
Published: (2023)
by: Zhou, Zhengping, et al.
Published: (2023)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
(How) Do Language Models Track State?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space
by: Ruscio, Valeria, et al.
Published: (2026)
by: Ruscio, Valeria, et al.
Published: (2026)
Large Language and Reasoning Models are Shallow Disjunctive Reasoners
by: Khalid, Irtaza, et al.
Published: (2025)
by: Khalid, Irtaza, et al.
Published: (2025)
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
by: Zhou, Andy, et al.
Published: (2023)
by: Zhou, Andy, et al.
Published: (2023)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Similar Items
-
Refusal in Language Models Is Mediated by a Single Direction
by: Arditi, Andy, et al.
Published: (2024) -
Real-Time Detection of Hallucinated Entities in Long-Form Generation
by: Obeso, Oscar, et al.
Published: (2025) -
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025) -
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025) -
Programming Refusal with Conditional Activation Steering
by: Lee, Bruce W., et al.
Published: (2024)