Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Jiahao, Luo, Haozheng, Hu, Jerry Yao-Chieh, Guo, Wenbo, Liu, Han, Xing, Xinyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decoupled Alignment for Robust Plug-and-Play Adaptation
von: Luo, Haozheng, et al.
Veröffentlicht: (2024)
von: Luo, Haozheng, et al.
Veröffentlicht: (2024)
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
von: Luo, Haozheng, et al.
Veröffentlicht: (2025)
von: Luo, Haozheng, et al.
Veröffentlicht: (2025)
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
BlockScan: Detecting Anomalies in Blockchain Transactions
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
von: Wang, Jerry, et al.
Veröffentlicht: (2025)
von: Wang, Jerry, et al.
Veröffentlicht: (2025)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
Attention Mechanism, Max-Affine Partition, and Universal Approximation
von: Liu, Hude, et al.
Veröffentlicht: (2025)
von: Liu, Hude, et al.
Veröffentlicht: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
von: Chen, Yu, et al.
Veröffentlicht: (2026)
von: Chen, Yu, et al.
Veröffentlicht: (2026)
In-Context Algorithm Emulation in Fixed-Weight Transformers
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
Universal Approximation with Softmax Attention
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models
von: Wu, Dennis, et al.
Veröffentlicht: (2024)
von: Wu, Dennis, et al.
Veröffentlicht: (2024)
Transformer Approximations from ReLUs
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2026)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action
von: Pan, Zhenyu, et al.
Veröffentlicht: (2024)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2024)
On Flow Matching KL Divergence
von: Su, Maojiang, et al.
Veröffentlicht: (2025)
von: Su, Maojiang, et al.
Veröffentlicht: (2025)
Synergistic Weak-Strong Collaboration by Aligning Preferences
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems
von: Peng, Dengyun, et al.
Veröffentlicht: (2025)
von: Peng, Dengyun, et al.
Veröffentlicht: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
von: Yu, Jiahao, et al.
Veröffentlicht: (2023)
von: Yu, Jiahao, et al.
Veröffentlicht: (2023)
A Survey on Explainable Deep Reinforcement Learning
von: Cheng, Zelei, et al.
Veröffentlicht: (2025)
von: Cheng, Zelei, et al.
Veröffentlicht: (2025)
Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
Question Classification with Deep Contextualized Transformer
von: Luo, Haozheng, et al.
Veröffentlicht: (2019)
von: Luo, Haozheng, et al.
Veröffentlicht: (2019)
Nonparametric Modern Hopfield Models
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
von: Ren, Xuancheng, et al.
Veröffentlicht: (2026)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
von: Mavi, John, et al.
Veröffentlicht: (2025)
von: Mavi, John, et al.
Veröffentlicht: (2025)
Differentially Private Kernel Density Estimation
von: Liu, Erzhi, et al.
Veröffentlicht: (2024)
von: Liu, Erzhi, et al.
Veröffentlicht: (2024)
Can Data-Driven Dynamics Reveal Hidden Physics? There Is A Need for Interpretable Neural Operators
von: Gao, Wenhan, et al.
Veröffentlicht: (2025)
von: Gao, Wenhan, et al.
Veröffentlicht: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
von: Yu, Jiahao, et al.
Veröffentlicht: (2025)
von: Yu, Jiahao, et al.
Veröffentlicht: (2025)
GPO: Learning from Critical Steps to Improve LLM Reasoning
von: Yu, Jiahao, et al.
Veröffentlicht: (2025)
von: Yu, Jiahao, et al.
Veröffentlicht: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
von: Weidener, Lukas, et al.
Veröffentlicht: (2026)
von: Weidener, Lukas, et al.
Veröffentlicht: (2026)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
von: Lin, Weiran, et al.
Veröffentlicht: (2024)
von: Lin, Weiran, et al.
Veröffentlicht: (2024)
ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
von: Noever, David, et al.
Veröffentlicht: (2025)
von: Noever, David, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Decoupled Alignment for Robust Plug-and-Play Adaptation
von: Luo, Haozheng, et al.
Veröffentlicht: (2024) -
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
von: Luo, Haozheng, et al.
Veröffentlicht: (2025) -
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
von: Luo, Haozheng, et al.
Veröffentlicht: (2026) -
BlockScan: Detecting Anomalies in Blockchain Transactions
von: Yu, Jiahao, et al.
Veröffentlicht: (2024) -
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)