LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ren, Xuancheng, Hu, Shijing, Lu, Zhihui, Huang, Jiangqi, Duan, Qiang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
par: Hu, Xulin, et autres
Publié: (2026)
par: Hu, Xulin, et autres
Publié: (2026)
Latent-space Attacks for Refusal Evasion in Language Models
par: Piras, Giorgio, et autres
Publié: (2026)
par: Piras, Giorgio, et autres
Publié: (2026)
Query Carefully: Detecting the Unanswerables in Text-to-SQL Tasks
par: Saxer, Jasmin, et autres
Publié: (2025)
par: Saxer, Jasmin, et autres
Publié: (2025)
LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
par: Shu, Huizhen, et autres
Publié: (2025)
par: Shu, Huizhen, et autres
Publié: (2025)
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
par: Dong, Mingwen, et autres
Publié: (2024)
par: Dong, Mingwen, et autres
Publié: (2024)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
par: Collu, Matteo Gioele, et autres
Publié: (2026)
par: Collu, Matteo Gioele, et autres
Publié: (2026)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
par: Alagharu, Rishab, et autres
Publié: (2026)
par: Alagharu, Rishab, et autres
Publié: (2026)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
par: Yuan, Youliang, et autres
Publié: (2024)
par: Yuan, Youliang, et autres
Publié: (2024)
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
par: Pattison, Cameron, et autres
Publié: (2026)
par: Pattison, Cameron, et autres
Publié: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
par: Muhamed, Aashiq, et autres
Publié: (2025)
par: Muhamed, Aashiq, et autres
Publié: (2025)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
par: Weidener, Lukas, et autres
Publié: (2026)
par: Weidener, Lukas, et autres
Publié: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
par: Si, Shengyun, et autres
Publié: (2025)
par: Si, Shengyun, et autres
Publié: (2025)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
par: García-Ferrero, Iker, et autres
Publié: (2025)
par: García-Ferrero, Iker, et autres
Publié: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
par: Pan, Wenbo, et autres
Publié: (2025)
par: Pan, Wenbo, et autres
Publié: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
par: von Recum, Alexander, et autres
Publié: (2024)
par: von Recum, Alexander, et autres
Publié: (2024)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
par: Cao, Lang
Publié: (2023)
par: Cao, Lang
Publié: (2023)
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
par: Yu, Jiahao, et autres
Publié: (2024)
par: Yu, Jiahao, et autres
Publié: (2024)
Where Do Reasoning Models Refuse?
par: Yamaguchi, Kureha, et autres
Publié: (2025)
par: Yamaguchi, Kureha, et autres
Publié: (2025)
Programming Refusal with Conditional Activation Steering
par: Lee, Bruce W., et autres
Publié: (2024)
par: Lee, Bruce W., et autres
Publié: (2024)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
par: Xie, Tinghao, et autres
Publié: (2024)
par: Xie, Tinghao, et autres
Publié: (2024)
Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation
par: Zhang, Wentao, et autres
Publié: (2026)
par: Zhang, Wentao, et autres
Publié: (2026)
Measuring and Eliminating Refusals in Military Large Language Models
par: FitzGerald, Jack, et autres
Publié: (2026)
par: FitzGerald, Jack, et autres
Publié: (2026)
Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment
par: Xue, Zhiyu, et autres
Publié: (2026)
par: Xue, Zhiyu, et autres
Publié: (2026)
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
par: Borah, Abhilekh, et autres
Publié: (2025)
par: Borah, Abhilekh, et autres
Publié: (2025)
Efficient Refusal Ablation in LLM through Optimal Transport
par: Nanfack, Geraldin, et autres
Publié: (2026)
par: Nanfack, Geraldin, et autres
Publié: (2026)
Linearly Decoding Refused Knowledge in Aligned Language Models
par: Shrivastava, Aryan, et autres
Publié: (2025)
par: Shrivastava, Aryan, et autres
Publié: (2025)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
par: Cui, Justin, et autres
Publié: (2024)
par: Cui, Justin, et autres
Publié: (2024)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
par: Liu, Zhenhua, et autres
Publié: (2024)
par: Liu, Zhenhua, et autres
Publié: (2024)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
par: Hu, Shijing, et autres
Publié: (2025)
par: Hu, Shijing, et autres
Publié: (2025)
Refusal as Silence: Gendered Disparities in Vision-Language Model Responses
par: Luo, Sha, et autres
Publié: (2024)
par: Luo, Sha, et autres
Publié: (2024)
Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement
par: Lu, Yuxiao, et autres
Publié: (2026)
par: Lu, Yuxiao, et autres
Publié: (2026)
LAECIPS: Large Vision Model Assisted Adaptive Edge-Cloud Collaboration for IoT-based Embodied Intelligence System
par: Hu, Shijing, et autres
Publié: (2024)
par: Hu, Shijing, et autres
Publié: (2024)
ICCU: In-Context Continual Unlearning via Pattern-Induced Refusal Rules
par: Pan, Ruihao, et autres
Publié: (2026)
par: Pan, Ruihao, et autres
Publié: (2026)
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
par: Rahimi, Eliron, et autres
Publié: (2026)
par: Rahimi, Eliron, et autres
Publié: (2026)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
par: Noever, David, et autres
Publié: (2025)
par: Noever, David, et autres
Publié: (2025)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
par: Hildebrandt, Fabian, et autres
Publié: (2025)
par: Hildebrandt, Fabian, et autres
Publié: (2025)
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
par: Zhang, Zhihao, et autres
Publié: (2026)
par: Zhang, Zhihao, et autres
Publié: (2026)
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
par: Wu, Tongxi, et autres
Publié: (2026)
par: Wu, Tongxi, et autres
Publié: (2026)
Documents similaires
-
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
par: Hu, Xulin, et autres
Publié: (2026) -
Latent-space Attacks for Refusal Evasion in Language Models
par: Piras, Giorgio, et autres
Publié: (2026) -
Query Carefully: Detecting the Unanswerables in Text-to-SQL Tasks
par: Saxer, Jasmin, et autres
Publié: (2025) -
LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
par: Shu, Huizhen, et autres
Publié: (2025) -
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
par: Dong, Mingwen, et autres
Publié: (2024)