Beyond Refusal: Probing the Limits of Agentic Self-Correction for Semantic Sensitive Information
Fuente:
arXiv
Guardado en:
| Autores principales: | Suleymanov, Umid, Rajabov, Zaur, Mirzazada, Emil, Kantarcioglu, Murat |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
por: Asgarov, Ali, et al.
Publicado: (2025)
por: Asgarov, Ali, et al.
Publicado: (2025)
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
por: Suleymanov, Umid, et al.
Publicado: (2026)
por: Suleymanov, Umid, et al.
Publicado: (2026)
SPRINT: Semi-supervised Prototypical Representation for Few-Shot Class-Incremental Tabular Learning
por: Suleymanov, Umid, et al.
Publicado: (2026)
por: Suleymanov, Umid, et al.
Publicado: (2026)
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
por: Xing, Sihao, et al.
Publicado: (2026)
por: Xing, Sihao, et al.
Publicado: (2026)
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
por: Isbarov, Jafar, et al.
Publicado: (2026)
por: Isbarov, Jafar, et al.
Publicado: (2026)
PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
por: Hussain, Khizar, et al.
Publicado: (2026)
por: Hussain, Khizar, et al.
Publicado: (2026)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
por: García-Ferrero, Iker, et al.
Publicado: (2025)
por: García-Ferrero, Iker, et al.
Publicado: (2025)
A Revealed Preference Framework for AI Alignment
por: Suleymanov, Elchin
Publicado: (2026)
por: Suleymanov, Elchin
Publicado: (2026)
FedDAG: Clustered Federated Learning via Global Data and Gradient Integration for Heterogeneous Environments
por: Pramanik, Anik, et al.
Publicado: (2026)
por: Pramanik, Anik, et al.
Publicado: (2026)
Using AI Uncertainty Quantification to Improve Human Decision-Making
por: Marusich, Laura R., et al.
Publicado: (2023)
por: Marusich, Laura R., et al.
Publicado: (2023)
Beyond Shortest Path: Agentic Vehicular Routing with Semantic Context
por: Braun, Carnot, et al.
Publicado: (2025)
por: Braun, Carnot, et al.
Publicado: (2025)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
por: Alagharu, Rishab, et al.
Publicado: (2026)
por: Alagharu, Rishab, et al.
Publicado: (2026)
Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
por: Hussain, Khizar, et al.
Publicado: (2026)
por: Hussain, Khizar, et al.
Publicado: (2026)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
por: Cao, Lang
Publicado: (2023)
por: Cao, Lang
Publicado: (2023)
Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
por: Yang, Kia-Jüng, et al.
Publicado: (2026)
por: Yang, Kia-Jüng, et al.
Publicado: (2026)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
por: Noever, David, et al.
Publicado: (2025)
por: Noever, David, et al.
Publicado: (2025)
Human Decision-Making with Persuasive and Narrative LLM Explanations
por: Marusich, Laura R., et al.
Publicado: (2026)
por: Marusich, Laura R., et al.
Publicado: (2026)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
por: Ren, Xuancheng, et al.
Publicado: (2026)
por: Ren, Xuancheng, et al.
Publicado: (2026)
Beyond the Attention Stability Boundary: Agentic Self-Synthesizing Reasoning Protocols
por: Shehata, Dahlia, et al.
Publicado: (2026)
por: Shehata, Dahlia, et al.
Publicado: (2026)
Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation
por: Zhang, Wentao, et al.
Publicado: (2026)
por: Zhang, Wentao, et al.
Publicado: (2026)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
por: Anonto, Riad Ahmed, et al.
Publicado: (2025)
por: Anonto, Riad Ahmed, et al.
Publicado: (2025)
VeriAct: Beyond Verifiability -- Agentic Synthesis of Correct and Complete Formal Specifications
por: Misu, Md Rakib Hossain, et al.
Publicado: (2026)
por: Misu, Md Rakib Hossain, et al.
Publicado: (2026)
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
por: Pattison, Cameron, et al.
Publicado: (2026)
por: Pattison, Cameron, et al.
Publicado: (2026)
Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection
por: Mukherjee, Kunal, et al.
Publicado: (2026)
por: Mukherjee, Kunal, et al.
Publicado: (2026)
Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward
por: Hu, Senkang, et al.
Publicado: (2026)
por: Hu, Senkang, et al.
Publicado: (2026)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
por: Prakash, Nirmalendu, et al.
Publicado: (2025)
por: Prakash, Nirmalendu, et al.
Publicado: (2025)
Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
por: Li, Zhuofeng, et al.
Publicado: (2026)
por: Li, Zhuofeng, et al.
Publicado: (2026)
Beyond Compliance: How AI Could Help Creative Writers by Refusing Them
por: Qin, Hua Xuan, et al.
Publicado: (2026)
por: Qin, Hua Xuan, et al.
Publicado: (2026)
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
por: Seah, Ee Wei, et al.
Publicado: (2026)
por: Seah, Ee Wei, et al.
Publicado: (2026)
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
por: Collu, Matteo Gioele, et al.
Publicado: (2026)
por: Collu, Matteo Gioele, et al.
Publicado: (2026)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
por: Weidener, Lukas, et al.
Publicado: (2026)
por: Weidener, Lukas, et al.
Publicado: (2026)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
por: Si, Shengyun, et al.
Publicado: (2025)
por: Si, Shengyun, et al.
Publicado: (2025)
Semantic Invariance in Agentic AI
por: de Zarzà, I., et al.
Publicado: (2026)
por: de Zarzà, I., et al.
Publicado: (2026)
Beyond Output Critique: Self-Correction via Task Distillation
por: Rahmani, Hossein A., et al.
Publicado: (2026)
por: Rahmani, Hossein A., et al.
Publicado: (2026)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
por: Yuan, Youliang, et al.
Publicado: (2024)
por: Yuan, Youliang, et al.
Publicado: (2024)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
por: Pan, Wenbo, et al.
Publicado: (2025)
por: Pan, Wenbo, et al.
Publicado: (2025)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
por: Guan, Zhong, et al.
Publicado: (2026)
por: Guan, Zhong, et al.
Publicado: (2026)
Beyond the 'Diff': Addressing Agentic Entropy in Agentic Software Development
por: Casserini, Matteo, et al.
Publicado: (2026)
por: Casserini, Matteo, et al.
Publicado: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
por: Muhamed, Aashiq, et al.
Publicado: (2025)
por: Muhamed, Aashiq, et al.
Publicado: (2025)
GeoSR: Cognitive-Agentic Framework for Probing Geospatial Knowledge Boundaries via Iterative Self-Refinement
por: Tang, Jinfan, et al.
Publicado: (2025)
por: Tang, Jinfan, et al.
Publicado: (2025)
Ejemplares similares
-
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
por: Asgarov, Ali, et al.
Publicado: (2025) -
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
por: Suleymanov, Umid, et al.
Publicado: (2026) -
SPRINT: Semi-supervised Prototypical Representation for Few-Shot Class-Incremental Tabular Learning
por: Suleymanov, Umid, et al.
Publicado: (2026) -
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
por: Xing, Sihao, et al.
Publicado: (2026) -
Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks
por: Isbarov, Jafar, et al.
Publicado: (2026)