Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Jiahao, Luo, Haozheng, Hu, Jerry Yao-Chieh, Guo, Wenbo, Liu, Han, Xing, Xinyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Decoupled Alignment for Robust Plug-and-Play Adaptation
di: Luo, Haozheng, et al.
Pubblicazione: (2024)
di: Luo, Haozheng, et al.
Pubblicazione: (2024)
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
di: Luo, Haozheng, et al.
Pubblicazione: (2025)
di: Luo, Haozheng, et al.
Pubblicazione: (2025)
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
BlockScan: Detecting Anomalies in Blockchain Transactions
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
di: Wang, Jerry, et al.
Pubblicazione: (2025)
di: Wang, Jerry, et al.
Pubblicazione: (2025)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Attention Mechanism, Max-Affine Partition, and Universal Approximation
di: Liu, Hude, et al.
Pubblicazione: (2025)
di: Liu, Hude, et al.
Pubblicazione: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
di: Yuan, Youliang, et al.
Pubblicazione: (2024)
di: Yuan, Youliang, et al.
Pubblicazione: (2024)
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
di: Chen, Yu, et al.
Pubblicazione: (2026)
di: Chen, Yu, et al.
Pubblicazione: (2026)
In-Context Algorithm Emulation in Fixed-Weight Transformers
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
Universal Approximation with Softmax Attention
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models
di: Wu, Dennis, et al.
Pubblicazione: (2024)
di: Wu, Dennis, et al.
Pubblicazione: (2024)
Transformer Approximations from ReLUs
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2026)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2025)
Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action
di: Pan, Zhenyu, et al.
Pubblicazione: (2024)
di: Pan, Zhenyu, et al.
Pubblicazione: (2024)
On Flow Matching KL Divergence
di: Su, Maojiang, et al.
Pubblicazione: (2025)
di: Su, Maojiang, et al.
Pubblicazione: (2025)
Synergistic Weak-Strong Collaboration by Aligning Preferences
di: Jiao, Yizhu, et al.
Pubblicazione: (2025)
di: Jiao, Yizhu, et al.
Pubblicazione: (2025)
Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems
di: Peng, Dengyun, et al.
Pubblicazione: (2025)
di: Peng, Dengyun, et al.
Pubblicazione: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
di: Shrivastava, Aryan, et al.
Pubblicazione: (2025)
di: Shrivastava, Aryan, et al.
Pubblicazione: (2025)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
di: Yu, Jiahao, et al.
Pubblicazione: (2023)
A Survey on Explainable Deep Reinforcement Learning
di: Cheng, Zelei, et al.
Pubblicazione: (2025)
di: Cheng, Zelei, et al.
Pubblicazione: (2025)
Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
di: Chen, Zijun, et al.
Pubblicazione: (2025)
di: Chen, Zijun, et al.
Pubblicazione: (2025)
Question Classification with Deep Contextualized Transformer
di: Luo, Haozheng, et al.
Pubblicazione: (2019)
di: Luo, Haozheng, et al.
Pubblicazione: (2019)
Nonparametric Modern Hopfield Models
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
di: Ren, Xuancheng, et al.
Pubblicazione: (2026)
di: Ren, Xuancheng, et al.
Pubblicazione: (2026)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
di: Mavi, John, et al.
Pubblicazione: (2025)
di: Mavi, John, et al.
Pubblicazione: (2025)
Differentially Private Kernel Density Estimation
di: Liu, Erzhi, et al.
Pubblicazione: (2024)
di: Liu, Erzhi, et al.
Pubblicazione: (2024)
Can Data-Driven Dynamics Reveal Hidden Physics? There Is A Need for Interpretable Neural Operators
di: Gao, Wenhan, et al.
Pubblicazione: (2025)
di: Gao, Wenhan, et al.
Pubblicazione: (2025)
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
di: Yu, Jiahao, et al.
Pubblicazione: (2025)
di: Yu, Jiahao, et al.
Pubblicazione: (2025)
GPO: Learning from Critical Steps to Improve LLM Reasoning
di: Yu, Jiahao, et al.
Pubblicazione: (2025)
di: Yu, Jiahao, et al.
Pubblicazione: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
di: Si, Shengyun, et al.
Pubblicazione: (2025)
di: Si, Shengyun, et al.
Pubblicazione: (2025)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
di: Weidener, Lukas, et al.
Pubblicazione: (2026)
di: Weidener, Lukas, et al.
Pubblicazione: (2026)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
di: Lin, Weiran, et al.
Pubblicazione: (2024)
di: Lin, Weiran, et al.
Pubblicazione: (2024)
ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
di: Zhang, Zhenliang, et al.
Pubblicazione: (2025)
di: Zhang, Zhenliang, et al.
Pubblicazione: (2025)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
di: Noever, David, et al.
Pubblicazione: (2025)
di: Noever, David, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Decoupled Alignment for Robust Plug-and-Play Adaptation
di: Luo, Haozheng, et al.
Pubblicazione: (2024) -
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
di: Luo, Haozheng, et al.
Pubblicazione: (2025) -
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
di: Luo, Haozheng, et al.
Pubblicazione: (2026) -
BlockScan: Detecting Anomalies in Blockchain Transactions
di: Yu, Jiahao, et al.
Pubblicazione: (2024) -
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
di: Pan, Wenbo, et al.
Pubblicazione: (2025)