Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Trivedi, Prashant, Chakraborty, Souradip, Reddy, Avinash, Aggarwal, Vaneet, Bedi, Amrit Singh, Atia, George K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
BalancedDPO: Adaptive Multi-Metric Alignment
by: Tamboli, Dipesh, et al.
Published: (2025)
by: Tamboli, Dipesh, et al.
Published: (2025)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)
by: Bai, Qinbo, et al.
Published: (2022)
Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
by: Gaur, Mudit, et al.
Published: (2024)
by: Gaur, Mudit, et al.
Published: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
by: Chehade, Mohamad, et al.
Published: (2025)
by: Chehade, Mohamad, et al.
Published: (2025)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
LANTERN: LLM-Augmented Neurosymbolic Transfer with Experience-Gated Reasoning Networks
by: Alinejad, Mahyar, et al.
Published: (2026)
by: Alinejad, Mahyar, et al.
Published: (2026)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
by: Reddy, Avinash, et al.
Published: (2026)
by: Reddy, Avinash, et al.
Published: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
PROPS: Progressively Private Self-alignment of Large Language Models
by: Teku, Noel, et al.
Published: (2025)
by: Teku, Noel, et al.
Published: (2025)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
by: Sun, Xingpeng, et al.
Published: (2024)
by: Sun, Xingpeng, et al.
Published: (2024)
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026)
by: Singh, Utsav, et al.
Published: (2026)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
Accelerating Quantum Reinforcement Learning with a Quantum Natural Policy Gradient Based Approach
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
$γ$-weakly $θ$-up-concavity: A Unified Framework for Non-Convex Optimization Beyond DR-Submodular and OSS Functions
by: Pedramfar, Mohammad, et al.
Published: (2026)
by: Pedramfar, Mohammad, et al.
Published: (2026)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
On The Global Convergence Of Online RLHF With Neural Parametrization
by: Gaur, Mudit, et al.
Published: (2024)
by: Gaur, Mudit, et al.
Published: (2024)
Sample Complexity Analysis for Constrained Bilevel Reinforcement Learning
by: Saxena, Naman, et al.
Published: (2026)
by: Saxena, Naman, et al.
Published: (2026)
On the Vulnerability of LLM/VLM-Controlled Robotics
by: Wu, Xiyang, et al.
Published: (2024)
by: Wu, Xiyang, et al.
Published: (2024)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
by: Lu, Yiyang, et al.
Published: (2025)
by: Lu, Yiyang, et al.
Published: (2025)
Joint Optimization of Multi-Objective Reinforcement Learning with Policy Gradient Based Algorithm
by: Bai, Qinbo, et al.
Published: (2021)
by: Bai, Qinbo, et al.
Published: (2021)
Sample-Efficient Constrained Reinforcement Learning with General Parameterization
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
by: Mondal, Washim Uddin, et al.
Published: (2023)
by: Mondal, Washim Uddin, et al.
Published: (2023)
Stronger Approximation Guarantees for Non-Monotone γ-Weakly DR-Submodular Maximization
by: Jadav, Hareshkumar, et al.
Published: (2026)
by: Jadav, Hareshkumar, et al.
Published: (2026)
A Unified Approach for Maximizing Continuous DR-submodular Functions
by: Pedramfar, Mohammad, et al.
Published: (2023)
by: Pedramfar, Mohammad, et al.
Published: (2023)
ECPv2: Fast, Efficient, and Scalable Global Optimization of Lipschitz Functions
by: Fourati, Fares, et al.
Published: (2025)
by: Fourati, Fares, et al.
Published: (2025)
Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback
by: Pedramfar, Mohammad, et al.
Published: (2023)
by: Pedramfar, Mohammad, et al.
Published: (2023)
Variational Offline Multi-agent Skill Discovery
by: Chen, Jiayu, et al.
Published: (2024)
by: Chen, Jiayu, et al.
Published: (2024)
Discrete State Diffusion Models: A Sample Complexity Perspective
by: Srikanth, Aadithya, et al.
Published: (2025)
by: Srikanth, Aadithya, et al.
Published: (2025)
Similar Items
-
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
by: Gaur, Mudit, et al.
Published: (2025) -
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
by: Barakat, Anas, et al.
Published: (2026) -
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
by: Gaur, Mudit, et al.
Published: (2025) -
BalancedDPO: Adaptive Multi-Metric Alignment
by: Tamboli, Dipesh, et al.
Published: (2025) -
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)