Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
Fuente:
arXiv
Saved in:
| Main Authors: | Pathmanathan, Pankayaraj, Huang, Furong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using Curiosity for an Even Representation of Tasks in Continual Offline Reinforcement Learning
by: Pathmanathan, Pankayaraj, et al.
Published: (2023)
by: Pathmanathan, Pankayaraj, et al.
Published: (2023)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
by: Pathmanathan, Pankayaraj, et al.
Published: (2025)
by: Pathmanathan, Pankayaraj, et al.
Published: (2025)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
Uncertainty-Aware Deep Learning Framework for Remaining Useful Life Prediction in Turbofan Engines with Learned Aleatoric Uncertainty
by: Sharma, Krishang
Published: (2025)
by: Sharma, Krishang
Published: (2025)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
Uncovering Memorization in Timeseries Imputation models: LBRM Membership Inference and its link to attribute Leakage
by: Taleb, Faiz, et al.
Published: (2026)
by: Taleb, Faiz, et al.
Published: (2026)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
by: Fei, Yu, et al.
Published: (2024)
by: Fei, Yu, et al.
Published: (2024)
Multi-step retrieval and reasoning improves radiology question answering with large language models
by: Wind, Sebastian, et al.
Published: (2025)
by: Wind, Sebastian, et al.
Published: (2025)
CNN-LSTM Hybrid Deep Learning Model for Remaining Useful Life Estimation
by: G, Muthukumar, et al.
Published: (2024)
by: G, Muthukumar, et al.
Published: (2024)
SAFLEX: Self-Adaptive Augmentation via Feature Label Extrapolation
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Unified Gradient-Based Machine Unlearning with Remain Geometry Enhancement
by: Huang, Zhehao, et al.
Published: (2024)
by: Huang, Zhehao, et al.
Published: (2024)
Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew
by: Serrà, Joan, et al.
Published: (2026)
by: Serrà, Joan, et al.
Published: (2026)
Test-time Correlation Alignment
by: You, Linjing, et al.
Published: (2025)
by: You, Linjing, et al.
Published: (2025)
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
LLM-FS-Agent: A Deliberative Role-based Large Language Model Architecture for Transparent Feature Selection
by: Bal-Ghaoui, Mohamed, et al.
Published: (2025)
by: Bal-Ghaoui, Mohamed, et al.
Published: (2025)
Agential AI for Integrated Continual Learning, Deliberative Behavior, and Comprehensible Models
by: Erden, Zeki Doruk, et al.
Published: (2025)
by: Erden, Zeki Doruk, et al.
Published: (2025)
Context information can be more important than reasoning for time series forecasting with a large language model
by: Yang, Janghoon
Published: (2025)
by: Yang, Janghoon
Published: (2025)
Boosting Sample Efficiency and Generalization in Multi-agent Reinforcement Learning via Equivariance
by: McClellan, Joshua, et al.
Published: (2024)
by: McClellan, Joshua, et al.
Published: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Deep Domain Adaptation for Turbofan Engine Remaining Useful Life Prediction: Methodologies, Evaluation and Future Trends
by: Wang, Yucheng, et al.
Published: (2025)
by: Wang, Yucheng, et al.
Published: (2025)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
Supervised Contrastive Learning based Dual-Mixer Model for Remaining Useful Life Prediction
by: Fu, En, et al.
Published: (2024)
by: Fu, En, et al.
Published: (2024)
Replacing thinking with tool usage enables reasoning in small language models
by: Rainone, Corrado, et al.
Published: (2025)
by: Rainone, Corrado, et al.
Published: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Spectral Greedy Coresets for Graph Neural Networks
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
by: Liu, Xiangyu, et al.
Published: (2023)
by: Liu, Xiangyu, et al.
Published: (2023)
Interactive dense pixel visualizations for time series and model attribution explanations
by: Schlegel, Udo, et al.
Published: (2024)
by: Schlegel, Udo, et al.
Published: (2024)
Learnable Chernoff Baselines for Inference-Time Alignment
by: Madhow, Sunil, et al.
Published: (2026)
by: Madhow, Sunil, et al.
Published: (2026)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
by: Huang, Audrey, et al.
Published: (2025)
by: Huang, Audrey, et al.
Published: (2025)
Reinforcing privacy reasoning in LLMs via normative simulacra from fiction
by: Franchi, Matt, et al.
Published: (2026)
by: Franchi, Matt, et al.
Published: (2026)
Training Deliberative Monitors for Black-Box Scheming Detection
by: Sinha, Aditya, et al.
Published: (2026)
by: Sinha, Aditya, et al.
Published: (2026)
Variational Inference via Smoothed Particle Hydrodynamics
by: Huang, Yongchao
Published: (2024)
by: Huang, Yongchao
Published: (2024)
Fixing confirmation bias in feature attribution methods via semantic match
by: Cinà, Giovanni, et al.
Published: (2023)
by: Cinà, Giovanni, et al.
Published: (2023)
Torch-Uncertainty: A Deep Learning Framework for Uncertainty Quantification
by: Lafage, Adrien, et al.
Published: (2025)
by: Lafage, Adrien, et al.
Published: (2025)
Automated Machine Learning for Remaining Useful Life Predictions
by: Zöller, Marc-André, et al.
Published: (2023)
by: Zöller, Marc-André, et al.
Published: (2023)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
Meta-Learning and Knowledge Discovery based Physics-Informed Neural Network for Remaining Useful Life Prediction
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Similar Items
-
Using Curiosity for an Even Representation of Tasks in Continual Offline Reinforcement Learning
by: Pathmanathan, Pankayaraj, et al.
Published: (2023) -
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
by: Pathmanathan, Pankayaraj, et al.
Published: (2024) -
Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
by: Pathmanathan, Pankayaraj, et al.
Published: (2025) -
Is poisoning a real threat to LLM alignment? Maybe more so than you think
by: Pathmanathan, Pankayaraj, et al.
Published: (2024) -
Uncertainty-Aware Deep Learning Framework for Remaining Useful Life Prediction in Turbofan Engines with Learned Aleatoric Uncertainty
by: Sharma, Krishang
Published: (2025)