Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Guobin, Zhao, Dongcheng, Tong, Haibo, Li, Jindong, Zhao, Feifei, Zeng, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
by: Han, Bing, et al.
Published: (2025)
by: Han, Bing, et al.
Published: (2025)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
by: Shen, Sicheng, et al.
Published: (2026)
by: Shen, Sicheng, et al.
Published: (2026)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
MVPBench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
by: Liang, Yao, et al.
Published: (2025)
by: Liang, Yao, et al.
Published: (2025)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Efficient LLM Safety Evaluation through Multi-Agent Debate
by: Lin, Dachuan, et al.
Published: (2025)
by: Lin, Dachuan, et al.
Published: (2025)
Developmental Plasticity-inspired Adaptive Pruning for Deep Spiking and Artificial Neural Networks
by: Han, Bing, et al.
Published: (2022)
by: Han, Bing, et al.
Published: (2022)
StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?
by: Shen, Guobin, et al.
Published: (2024)
by: Shen, Guobin, et al.
Published: (2024)
Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind
by: Tong, Haibo, et al.
Published: (2024)
by: Tong, Haibo, et al.
Published: (2024)
Evolving Efficient Genetic Encoding for Deep Spiking Neural Networks
by: Pan, Wenxuan, et al.
Published: (2024)
by: Pan, Wenxuan, et al.
Published: (2024)
Internalizing Safety Understanding in Large Reasoning Models via Verification
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
Trust & Safety of LLMs and LLMs in Trust & Safety
by: You, Doohee, et al.
Published: (2024)
by: You, Doohee, et al.
Published: (2024)
$SpikePack$: Enhanced Information Flow in Spiking Neural Networks with High Hardware Compatibility
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
Learning the Plasticity: Plasticity-Driven Learning Framework in Spiking Neural Networks
by: Shen, Guobin, et al.
Published: (2023)
by: Shen, Guobin, et al.
Published: (2023)
Brain-inspired and Self-based Artificial Intelligence
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
Building Altruistic and Moral AI Agent with Brain-inspired Emotional Empathy Mechanisms
by: Zhao, Feifei, et al.
Published: (2024)
by: Zhao, Feifei, et al.
Published: (2024)
Multi-compartment Neuron and Population Encoding Powered Spiking Neural Network for Deep Distributional Reinforcement Learning
by: Sun, Yinqian, et al.
Published: (2023)
by: Sun, Yinqian, et al.
Published: (2023)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
by: Li, Tenglong, et al.
Published: (2025)
by: Li, Tenglong, et al.
Published: (2025)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
by: Li, Tenglong, et al.
Published: (2024)
by: Li, Tenglong, et al.
Published: (2024)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
by: Li, Tenglong, et al.
Published: (2026)
by: Li, Tenglong, et al.
Published: (2026)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
by: Li, Jindong, et al.
Published: (2024)
by: Li, Jindong, et al.
Published: (2024)
Super Co-alignment of Human and AI for Sustainable Symbiotic Society
by: Zeng, Yi, et al.
Published: (2025)
by: Zeng, Yi, et al.
Published: (2025)
CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
Continual Learning of Multiple Cognitive Functions with Brain-inspired Temporal Development Mechanism
by: Han, Bing, et al.
Published: (2025)
by: Han, Bing, et al.
Published: (2025)
Compass-Thinker-7B Technical Report
by: Zeng, Anxiang, et al.
Published: (2025)
by: Zeng, Anxiang, et al.
Published: (2025)
Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks
by: Wang, Jihang, et al.
Published: (2025)
by: Wang, Jihang, et al.
Published: (2025)
TIM: An Efficient Temporal Interaction Module for Spiking Transformer
by: Shen, Sicheng, et al.
Published: (2024)
by: Shen, Sicheng, et al.
Published: (2024)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
by: Zhang, Dongcheng, et al.
Published: (2026)
by: Zhang, Dongcheng, et al.
Published: (2026)
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
by: Shen, Guobin, et al.
Published: (2024)
by: Shen, Guobin, et al.
Published: (2024)
A Brain-inspired Memory Transformation based Differentiable Neural Computer for Reasoning-based Question Answering
by: Liang, Yao, et al.
Published: (2023)
by: Liang, Yao, et al.
Published: (2023)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
by: Li, Xiaohu, et al.
Published: (2025)
by: Li, Xiaohu, et al.
Published: (2025)
Adaptive Reorganization of Neural Pathways for Continual Learning with Spiking Neural Networks
by: Han, Bing, et al.
Published: (2023)
by: Han, Bing, et al.
Published: (2023)
TEFormer: Structured Bidirectional Temporal Enhancement Modeling in Spiking Transformers
by: Shen, Sicheng, et al.
Published: (2026)
by: Shen, Sicheng, et al.
Published: (2026)
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
Similar Items
-
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025) -
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
by: Han, Bing, et al.
Published: (2025) -
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
by: Shen, Sicheng, et al.
Published: (2026) -
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
by: Shen, Guobin, et al.
Published: (2026) -
MVPBench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
by: Liang, Yao, et al.
Published: (2025)