MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhai, Kevin, Singh, Utsav, Thatipelli, Anirudh, Chakraborty, Souradip, Sahu, Anit Kumar, Huang, Furong, Bedi, Amrit Singh, Shah, Mubarak |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
por: Beetham, James, et al.
Publicado: (2024)
por: Beetham, James, et al.
Publicado: (2024)
RL with Learnable Textual Feedback: A Bilevel Approach
por: Singh, Utsav, et al.
Publicado: (2026)
por: Singh, Utsav, et al.
Publicado: (2026)
Towards Realistic Mechanisms That Incentivize Federated Participation and Contribution
por: Bornstein, Marco, et al.
Publicado: (2023)
por: Bornstein, Marco, et al.
Publicado: (2023)
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
por: Singh, Utsav, et al.
Publicado: (2024)
por: Singh, Utsav, et al.
Publicado: (2024)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
por: Lee, Jihoon, et al.
Publicado: (2025)
por: Lee, Jihoon, et al.
Publicado: (2025)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
por: Agrawal, Aakriti, et al.
Publicado: (2025)
por: Agrawal, Aakriti, et al.
Publicado: (2025)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
por: Chakraborty, Souradip, et al.
Publicado: (2023)
por: Chakraborty, Souradip, et al.
Publicado: (2023)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
por: Chakraborty, Souradip, et al.
Publicado: (2023)
por: Chakraborty, Souradip, et al.
Publicado: (2023)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2026)
Transfer Q Star: Principled Decoding for LLM Alignment
por: Chakraborty, Souradip, et al.
Publicado: (2024)
por: Chakraborty, Souradip, et al.
Publicado: (2024)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
por: Barakat, Anas, et al.
Publicado: (2026)
por: Barakat, Anas, et al.
Publicado: (2026)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
por: Chehade, Mohamad, et al.
Publicado: (2025)
por: Chehade, Mohamad, et al.
Publicado: (2025)
MaxMin-RLHF: Alignment with Diverse Human Preferences
por: Chakraborty, Souradip, et al.
Publicado: (2024)
por: Chakraborty, Souradip, et al.
Publicado: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
por: Ghosal, Soumya Suvra, et al.
Publicado: (2024)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2024)
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
por: Singh, Utsav, et al.
Publicado: (2024)
por: Singh, Utsav, et al.
Publicado: (2024)
BalancedDPO: Adaptive Multi-Metric Alignment
por: Tamboli, Dipesh, et al.
Publicado: (2025)
por: Tamboli, Dipesh, et al.
Publicado: (2025)
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
por: Yousaf, Adeel, et al.
Publicado: (2025)
por: Yousaf, Adeel, et al.
Publicado: (2025)
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
por: Veeramachaneni, Keerthi, et al.
Publicado: (2025)
por: Veeramachaneni, Keerthi, et al.
Publicado: (2025)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
por: Trivedi, Prashant, et al.
Publicado: (2025)
por: Trivedi, Prashant, et al.
Publicado: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
por: Barakat, Anas, et al.
Publicado: (2024)
por: Barakat, Anas, et al.
Publicado: (2024)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
por: Sun, Xingpeng, et al.
Publicado: (2024)
por: Sun, Xingpeng, et al.
Publicado: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
por: Patel, Bhrij, et al.
Publicado: (2024)
por: Patel, Bhrij, et al.
Publicado: (2024)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
por: Ding, Mucong, et al.
Publicado: (2024)
por: Ding, Mucong, et al.
Publicado: (2024)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
por: Gaur, Mudit, et al.
Publicado: (2025)
por: Gaur, Mudit, et al.
Publicado: (2025)
PROPS: Progressively Private Self-alignment of Large Language Models
por: Teku, Noel, et al.
Publicado: (2025)
por: Teku, Noel, et al.
Publicado: (2025)
FACT or Fiction: Can Truthful Mechanisms Eliminate Federated Free Riding?
por: Bornstein, Marco, et al.
Publicado: (2024)
por: Bornstein, Marco, et al.
Publicado: (2024)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
por: Singh, Anukriti, et al.
Publicado: (2025)
por: Singh, Anukriti, et al.
Publicado: (2025)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
por: Ghosal, Soumya Suvra, et al.
Publicado: (2025)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2025)
AI Cap-and-Trade: Efficiency Incentives for Accessibility and Sustainability
por: Bornstein, Marco, et al.
Publicado: (2026)
por: Bornstein, Marco, et al.
Publicado: (2026)
A lift for input-convex neural network training
por: Siahkoohi, Ali, et al.
Publicado: (2026)
por: Siahkoohi, Ali, et al.
Publicado: (2026)
Hypernetwork-based approach for grid-independent functional data clustering
por: Thatipelli, Anirudh, et al.
Publicado: (2026)
por: Thatipelli, Anirudh, et al.
Publicado: (2026)
PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling
por: Singh, Utsav, et al.
Publicado: (2024)
por: Singh, Utsav, et al.
Publicado: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Dystruct: Dynamically Structured Diffusion Language Model Decoding via Bayesian Inference
por: Sun, Bian, et al.
Publicado: (2026)
por: Sun, Bian, et al.
Publicado: (2026)
On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows
por: Chakraborty, Souradip, et al.
Publicado: (2025)
por: Chakraborty, Souradip, et al.
Publicado: (2025)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
por: Bai, Qinbo, et al.
Publicado: (2022)
por: Bai, Qinbo, et al.
Publicado: (2022)
Auction-Based Regulation for Artificial Intelligence
por: Bornstein, Marco, et al.
Publicado: (2024)
por: Bornstein, Marco, et al.
Publicado: (2024)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
por: Deng, Wenlong, et al.
Publicado: (2026)
por: Deng, Wenlong, et al.
Publicado: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
On the Vulnerability of LLM/VLM-Controlled Robotics
por: Wu, Xiyang, et al.
Publicado: (2024)
por: Wu, Xiyang, et al.
Publicado: (2024)
Ejemplares similares
-
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
por: Beetham, James, et al.
Publicado: (2024) -
RL with Learnable Textual Feedback: A Bilevel Approach
por: Singh, Utsav, et al.
Publicado: (2026) -
Towards Realistic Mechanisms That Incentivize Federated Participation and Contribution
por: Bornstein, Marco, et al.
Publicado: (2023) -
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
por: Singh, Utsav, et al.
Publicado: (2024) -
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
por: Lee, Jihoon, et al.
Publicado: (2025)