AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Weixiang, Guo, Jiahe, Hu, Yulin, Deng, Yang, Zhang, An, Sui, Xingyu, Han, Xinyang, Zhao, Yanyan, Qin, Bing, Chua, Tat-Seng, Liu, Ting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
von: Sui, Xingyu, et al.
Veröffentlicht: (2026)
von: Sui, Xingyu, et al.
Veröffentlicht: (2026)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
MPO: Multilingual Safety Alignment via Reward Gap Optimization
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
von: Hu, Yulin, et al.
Veröffentlicht: (2026)
von: Hu, Yulin, et al.
Veröffentlicht: (2026)
Chain of Strategy Optimization Makes Large Language Models Better Emotional Supporter
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
Lens: Rethinking Multilingual Enhancement for Large Language Models
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
von: Guo, Jiahe, et al.
Veröffentlicht: (2026)
von: Guo, Jiahe, et al.
Veröffentlicht: (2026)
Understanding Multilingualism in Mixture-of-Experts LLMs: Routing Mechanism, Expert Specialization, and Layerwise Steering
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
Learning to Learn from Multimodal Experience
von: Sui, Xingyu, et al.
Veröffentlicht: (2026)
von: Sui, Xingyu, et al.
Veröffentlicht: (2026)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
SteerX: Disentangled Steering for LLM Personalization
von: Zhao, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaoyan, et al.
Veröffentlicht: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction
von: Guo, Jiahe, et al.
Veröffentlicht: (2026)
von: Guo, Jiahe, et al.
Veröffentlicht: (2026)
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
von: Zheng, Jingnan, et al.
Veröffentlicht: (2025)
von: Zheng, Jingnan, et al.
Veröffentlicht: (2025)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents
von: Fu, Xing, et al.
Veröffentlicht: (2026)
von: Fu, Xing, et al.
Veröffentlicht: (2026)
Both Matter: Enhancing the Emotional Intelligence of Large Language Models without Compromising the General Intelligence
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Self-Guard: Defending Large Reasoning Models via enhanced self-reflection
von: Zheng, Jingnan, et al.
Veröffentlicht: (2026)
von: Zheng, Jingnan, et al.
Veröffentlicht: (2026)
Adaptive Probe-based Steering for Robust LLM Jailbreaking
von: Chen, Junxi, et al.
Veröffentlicht: (2026)
von: Chen, Junxi, et al.
Veröffentlicht: (2026)
A Federated Framework for LLM-based Recommendation
von: Zhao, Jujia, et al.
Veröffentlicht: (2024)
von: Zhao, Jujia, et al.
Veröffentlicht: (2024)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
von: Deng, Yang, et al.
Veröffentlicht: (2023)
von: Deng, Yang, et al.
Veröffentlicht: (2023)
On Reasoning Strength Planning in Large Reasoning Models
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations
von: Qin, Peixin, et al.
Veröffentlicht: (2024)
von: Qin, Peixin, et al.
Veröffentlicht: (2024)
SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language Models
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
von: Cui, Chenhang, et al.
Veröffentlicht: (2024)
von: Cui, Chenhang, et al.
Veröffentlicht: (2024)
Enhancing Spectral Graph Neural Networks with LLM-Predicted Homophily
von: Lu, Kangkang, et al.
Veröffentlicht: (2025)
von: Lu, Kangkang, et al.
Veröffentlicht: (2025)
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2026)
Personalized Text Generation with Contrastive Activation Steering
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025) -
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025) -
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024) -
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
von: Sui, Xingyu, et al.
Veröffentlicht: (2026) -
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)