Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chehade, Mohamad, Ghosal, Soumya Suvra, Chakraborty, Souradip, Reddy, Avinash, Manocha, Dinesh, Zhu, Hao, Bedi, Amrit Singh |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
par: Ghosal, Soumya Suvra, et autres
Publié: (2026)
par: Ghosal, Soumya Suvra, et autres
Publié: (2026)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
par: Ghosal, Soumya Suvra, et autres
Publié: (2025)
par: Ghosal, Soumya Suvra, et autres
Publié: (2025)
Transfer Q Star: Principled Decoding for LLM Alignment
par: Chakraborty, Souradip, et autres
Publié: (2024)
par: Chakraborty, Souradip, et autres
Publié: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
par: Patel, Bhrij, et autres
Publié: (2024)
par: Patel, Bhrij, et autres
Publié: (2024)
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
par: Chakraborty, Souradip, et autres
Publié: (2024)
par: Chakraborty, Souradip, et autres
Publié: (2024)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
par: Chakraborty, Souradip, et autres
Publié: (2025)
par: Chakraborty, Souradip, et autres
Publié: (2025)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
par: Beetham, James, et autres
Publié: (2024)
par: Beetham, James, et autres
Publié: (2024)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
par: Reddy, Avinash, et autres
Publié: (2026)
par: Reddy, Avinash, et autres
Publié: (2026)
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
par: Ghosal, Soumya Suvra, et autres
Publié: (2025)
par: Ghosal, Soumya Suvra, et autres
Publié: (2025)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
par: Trivedi, Prashant, et autres
Publié: (2025)
par: Trivedi, Prashant, et autres
Publié: (2025)
Multi-LLM QA with Embodied Exploration
par: Patel, Bhrij, et autres
Publié: (2024)
par: Patel, Bhrij, et autres
Publié: (2024)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
par: Dogra, Atharvan, et autres
Publié: (2025)
par: Dogra, Atharvan, et autres
Publié: (2025)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
par: Ding, Mucong, et autres
Publié: (2024)
par: Ding, Mucong, et autres
Publié: (2024)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
par: Sun, Xingpeng, et autres
Publié: (2024)
par: Sun, Xingpeng, et autres
Publié: (2024)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
par: Barakat, Anas, et autres
Publié: (2026)
par: Barakat, Anas, et autres
Publié: (2026)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
par: Chakraborty, Souradip, et autres
Publié: (2023)
par: Chakraborty, Souradip, et autres
Publié: (2023)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
par: Chakraborty, Souradip, et autres
Publié: (2023)
par: Chakraborty, Souradip, et autres
Publié: (2023)
BalancedDPO: Adaptive Multi-Metric Alignment
par: Tamboli, Dipesh, et autres
Publié: (2025)
par: Tamboli, Dipesh, et autres
Publié: (2025)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
par: Singh, Vaibhav, et autres
Publié: (2025)
par: Singh, Vaibhav, et autres
Publié: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
par: Barakat, Anas, et autres
Publié: (2024)
par: Barakat, Anas, et autres
Publié: (2024)
On the Vulnerability of LLM/VLM-Controlled Robotics
par: Wu, Xiyang, et autres
Publié: (2024)
par: Wu, Xiyang, et autres
Publié: (2024)
Fine-Tuning LLMs to Generate Economical and Reliable Actions for the Power Grid
par: Chehade, Mohamad, et autres
Publié: (2026)
par: Chehade, Mohamad, et autres
Publié: (2026)
TRAM: Test-Time Risk Adaptation with Mixture of Agents
par: Chehade, Mohamad Fares El Hajj, et autres
Publié: (2024)
par: Chehade, Mohamad Fares El Hajj, et autres
Publié: (2024)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
par: Agrawal, Aakriti, et autres
Publié: (2025)
par: Agrawal, Aakriti, et autres
Publié: (2025)
RL with Learnable Textual Feedback: A Bilevel Approach
par: Singh, Utsav, et autres
Publié: (2026)
par: Singh, Utsav, et autres
Publié: (2026)
DMCA: Dense Multi-agent Navigation using Attention and Communication
par: Arul, Senthil Hariharan, et autres
Publié: (2022)
par: Arul, Senthil Hariharan, et autres
Publié: (2022)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
par: Lee, Jihoon, et autres
Publié: (2025)
par: Lee, Jihoon, et autres
Publié: (2025)
PROPS: Progressively Private Self-alignment of Large Language Models
par: Teku, Noel, et autres
Publié: (2025)
par: Teku, Noel, et autres
Publié: (2025)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
par: Zhai, Kevin, et autres
Publié: (2025)
par: Zhai, Kevin, et autres
Publié: (2025)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
par: Singh, Anukriti, et autres
Publié: (2025)
par: Singh, Anukriti, et autres
Publié: (2025)
Enhancing Deep Neural Network Reliability with Refinement and Calibration
par: Hebbalaguppe, Ramya, et autres
Publié: (2026)
par: Hebbalaguppe, Ramya, et autres
Publié: (2026)
On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows
par: Chakraborty, Souradip, et autres
Publié: (2025)
par: Chakraborty, Souradip, et autres
Publié: (2025)
Personalized Embodied Navigation for Portable Object Finding
par: Dorbala, Vishnu Sashank, et autres
Publié: (2024)
par: Dorbala, Vishnu Sashank, et autres
Publié: (2024)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
par: Fei, Yu, et autres
Publié: (2024)
par: Fei, Yu, et autres
Publié: (2024)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
par: Cheng, Stephen, et autres
Publié: (2026)
par: Cheng, Stephen, et autres
Publié: (2026)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
par: Gaur, Mudit, et autres
Publié: (2025)
par: Gaur, Mudit, et autres
Publié: (2025)
Quantifying and Mitigating Premature Closure in Frontier LLMs
par: Handler, Rebecca, et autres
Publié: (2026)
par: Handler, Rebecca, et autres
Publié: (2026)
Documents similaires
-
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
par: Ghosal, Soumya Suvra, et autres
Publié: (2026) -
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
par: Ghosal, Soumya Suvra, et autres
Publié: (2025) -
Transfer Q Star: Principled Decoding for LLM Alignment
par: Chakraborty, Souradip, et autres
Publié: (2024) -
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
par: Ghosal, Soumya Suvra, et autres
Publié: (2024) -
Code Comprehension then Auditing for Unsupervised LLM Evaluation
par: Patel, Bhrij, et autres
Publié: (2024)