Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Ghosal, Soumya Suvra, Chakraborty, Souradip, Singh, Vaibhav, Guan, Tianrui, Wang, Mengdi, Velasquez, Alvaro, Beirami, Ahmad, Huang, Furong, Manocha, Dinesh, Bedi, Amrit Singh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
by: Chehade, Mohamad, et al.
Published: (2025)
by: Chehade, Mohamad, et al.
Published: (2025)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
by: Beetham, James, et al.
Published: (2024)
by: Beetham, James, et al.
Published: (2024)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
by: Dogra, Atharvan, et al.
Published: (2025)
by: Dogra, Atharvan, et al.
Published: (2025)
On the Vulnerability of LLM/VLM-Controlled Robotics
by: Wu, Xiyang, et al.
Published: (2024)
by: Wu, Xiyang, et al.
Published: (2024)
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026)
by: Singh, Utsav, et al.
Published: (2026)
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
DMCA: Dense Multi-agent Navigation using Attention and Communication
by: Arul, Senthil Hariharan, et al.
Published: (2022)
by: Arul, Senthil Hariharan, et al.
Published: (2022)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
by: Zhai, Kevin, et al.
Published: (2025)
by: Zhai, Kevin, et al.
Published: (2025)
BalancedDPO: Adaptive Multi-Metric Alignment
by: Tamboli, Dipesh, et al.
Published: (2025)
by: Tamboli, Dipesh, et al.
Published: (2025)
On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
Personalized Embodied Navigation for Portable Object Finding
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
by: Dorbala, Vishnu Sashank, et al.
Published: (2024)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
by: Trivedi, Prashant, et al.
Published: (2025)
by: Trivedi, Prashant, et al.
Published: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
by: Sun, Xingpeng, et al.
Published: (2024)
by: Sun, Xingpeng, et al.
Published: (2024)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
PROPS: Progressively Private Self-alignment of Large Language Models
by: Teku, Noel, et al.
Published: (2025)
by: Teku, Noel, et al.
Published: (2025)
FACT or Fiction: Can Truthful Mechanisms Eliminate Federated Free Riding?
by: Bornstein, Marco, et al.
Published: (2024)
by: Bornstein, Marco, et al.
Published: (2024)
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
by: Wu, Xiyang, et al.
Published: (2026)
by: Wu, Xiyang, et al.
Published: (2026)
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
AI Cap-and-Trade: Efficiency Incentives for Accessibility and Sustainability
by: Bornstein, Marco, et al.
Published: (2026)
by: Bornstein, Marco, et al.
Published: (2026)
Towards Realistic Mechanisms That Incentivize Federated Participation and Contribution
by: Bornstein, Marco, et al.
Published: (2023)
by: Bornstein, Marco, et al.
Published: (2023)
Paired-CSLiDAR: Height-Stratified Registration for Cross-Source Aerial-Ground LiDAR Pose Refinement
by: Hoover, Montana, et al.
Published: (2026)
by: Hoover, Montana, et al.
Published: (2026)
Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
Enhancing Deep Neural Network Reliability with Refinement and Calibration
by: Hebbalaguppe, Ramya, et al.
Published: (2026)
by: Hebbalaguppe, Ramya, et al.
Published: (2026)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)
by: Bai, Qinbo, et al.
Published: (2022)
Similar Items
-
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026) -
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024) -
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
by: Chehade, Mohamad, et al.
Published: (2025) -
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
by: Ghosal, Soumya Suvra, et al.
Published: (2025) -
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
by: Beetham, James, et al.
Published: (2024)