Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Guobin, Zhao, Dongcheng, Tong, Haibo, Li, Jindong, Zhao, Feifei, Zeng, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
MVPBench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
von: Liang, Yao, et al.
Veröffentlicht: (2025)
von: Liang, Yao, et al.
Veröffentlicht: (2025)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Efficient LLM Safety Evaluation through Multi-Agent Debate
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
von: Lin, Dachuan, et al.
Veröffentlicht: (2025)
Developmental Plasticity-inspired Adaptive Pruning for Deep Spiking and Artificial Neural Networks
von: Han, Bing, et al.
Veröffentlicht: (2022)
von: Han, Bing, et al.
Veröffentlicht: (2022)
StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind
von: Tong, Haibo, et al.
Veröffentlicht: (2024)
von: Tong, Haibo, et al.
Veröffentlicht: (2024)
Evolving Efficient Genetic Encoding for Deep Spiking Neural Networks
von: Pan, Wenxuan, et al.
Veröffentlicht: (2024)
von: Pan, Wenxuan, et al.
Veröffentlicht: (2024)
Internalizing Safety Understanding in Large Reasoning Models via Verification
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
Trust & Safety of LLMs and LLMs in Trust & Safety
von: You, Doohee, et al.
Veröffentlicht: (2024)
von: You, Doohee, et al.
Veröffentlicht: (2024)
$SpikePack$: Enhanced Information Flow in Spiking Neural Networks with High Hardware Compatibility
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
Learning the Plasticity: Plasticity-Driven Learning Framework in Spiking Neural Networks
von: Shen, Guobin, et al.
Veröffentlicht: (2023)
von: Shen, Guobin, et al.
Veröffentlicht: (2023)
Brain-inspired and Self-based Artificial Intelligence
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
von: Yao, Jing, et al.
Veröffentlicht: (2025)
von: Yao, Jing, et al.
Veröffentlicht: (2025)
Building Altruistic and Moral AI Agent with Brain-inspired Emotional Empathy Mechanisms
von: Zhao, Feifei, et al.
Veröffentlicht: (2024)
von: Zhao, Feifei, et al.
Veröffentlicht: (2024)
Multi-compartment Neuron and Population Encoding Powered Spiking Neural Network for Deep Distributional Reinforcement Learning
von: Sun, Yinqian, et al.
Veröffentlicht: (2023)
von: Sun, Yinqian, et al.
Veröffentlicht: (2023)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025)
von: Li, Jindong, et al.
Veröffentlicht: (2025)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
von: Li, Jindong, et al.
Veröffentlicht: (2024)
von: Li, Jindong, et al.
Veröffentlicht: (2024)
Super Co-alignment of Human and AI for Sustainable Symbiotic Society
von: Zeng, Yi, et al.
Veröffentlicht: (2025)
von: Zeng, Yi, et al.
Veröffentlicht: (2025)
CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
Continual Learning of Multiple Cognitive Functions with Brain-inspired Temporal Development Mechanism
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
Compass-Thinker-7B Technical Report
von: Zeng, Anxiang, et al.
Veröffentlicht: (2025)
von: Zeng, Anxiang, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks
von: Wang, Jihang, et al.
Veröffentlicht: (2025)
von: Wang, Jihang, et al.
Veröffentlicht: (2025)
TIM: An Efficient Temporal Interaction Module for Spiking Transformer
von: Shen, Sicheng, et al.
Veröffentlicht: (2024)
von: Shen, Sicheng, et al.
Veröffentlicht: (2024)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
von: Tong, Haibo, et al.
Veröffentlicht: (2026)
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
von: Zhang, Dongcheng, et al.
Veröffentlicht: (2026)
von: Zhang, Dongcheng, et al.
Veröffentlicht: (2026)
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
A Brain-inspired Memory Transformation based Differentiable Neural Computer for Reasoning-based Question Answering
von: Liang, Yao, et al.
Veröffentlicht: (2023)
von: Liang, Yao, et al.
Veröffentlicht: (2023)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
Adaptive Reorganization of Neural Pathways for Continual Learning with Spiking Neural Networks
von: Han, Bing, et al.
Veröffentlicht: (2023)
von: Han, Bing, et al.
Veröffentlicht: (2023)
TEFormer: Structured Bidirectional Temporal Enhancement Modeling in Spiking Transformers
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
von: Ueareeworakul, Pakorn, et al.
Veröffentlicht: (2025)
von: Ueareeworakul, Pakorn, et al.
Veröffentlicht: (2025)
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025)
von: Li, Jindong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025) -
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
von: Han, Bing, et al.
Veröffentlicht: (2025) -
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
von: Shen, Sicheng, et al.
Veröffentlicht: (2026) -
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026) -
MVPBench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
von: Liang, Yao, et al.
Veröffentlicht: (2025)