One-Shot Safety Alignment for Large Language Models via Optimal Dualization
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Xinmeng, Li, Shuo, Dobriban, Edgar, Bastani, Osbert, Hassani, Hamed, Ding, Dongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty in Language Models: Assessment through Rank-Calibration
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Alignment of large language models with constrained learning
by: Zhang, Botong, et al.
Published: (2025)
by: Zhang, Botong, et al.
Published: (2025)
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
by: Mitra, Aritra, et al.
Published: (2023)
by: Mitra, Aritra, et al.
Published: (2023)
Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
by: Huang, Xinmeng, et al.
Published: (2023)
by: Huang, Xinmeng, et al.
Published: (2023)
Reward Collapse in Aligning Large Language Models
by: Song, Ziang, et al.
Published: (2023)
by: Song, Ziang, et al.
Published: (2023)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
Leveraging Large Language Models for Solving Rare MIP Challenges
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
Understanding the Influence of Digraphs on Decentralized Optimization: Effective Metrics, Lower Bound, and Optimal Algorithm
by: Liang, Liyuan, et al.
Published: (2023)
by: Liang, Liyuan, et al.
Published: (2023)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
by: Li, Shuo, et al.
Published: (2023)
by: Li, Shuo, et al.
Published: (2023)
Deterministic Policy Gradient Primal-Dual Methods for Continuous-Space Constrained MDPs
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Hyperparameter Optimization for Large Language Model Instruction-Tuning
by: Tribes, Christophe, et al.
Published: (2023)
by: Tribes, Christophe, et al.
Published: (2023)
COS-DPO: Conditioned One-Shot Multi-Objective Fine-Tuning Framework
by: Ren, Yinuo, et al.
Published: (2024)
by: Ren, Yinuo, et al.
Published: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Watermarking Language Models with Error Correcting Codes
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
Exploiting Edited Large Language Models as General Scientific Optimizers
by: Lv, Qitan, et al.
Published: (2025)
by: Lv, Qitan, et al.
Published: (2025)
Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning
by: Qian, Tianhao, et al.
Published: (2026)
by: Qian, Tianhao, et al.
Published: (2026)
Variational Learning is Effective for Large Deep Networks
by: Shen, Yuesong, et al.
Published: (2024)
by: Shen, Yuesong, et al.
Published: (2024)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2022)
by: Ding, Dongsheng, et al.
Published: (2022)
Large-scale Online Ridesharing: The Effect of Assignment Optimality on System Performance
by: Fiedler, David, et al.
Published: (2023)
by: Fiedler, David, et al.
Published: (2023)
Advances and Challenges in Semantic Textual Similarity: A Comprehensive Survey
by: Kumar, Lokendra, et al.
Published: (2025)
by: Kumar, Lokendra, et al.
Published: (2025)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
OR-R1: Automating Modeling and Solving of Operations Research Optimization Problem via Test-Time Reinforcement Learning
by: Ding, Zezhen, et al.
Published: (2025)
by: Ding, Zezhen, et al.
Published: (2025)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Large Language Model-Assisted Planning of Electric Vehicle Charging Infrastructure with Real-World Case Study
by: Zheng, Xinda, et al.
Published: (2025)
by: Zheng, Xinda, et al.
Published: (2025)
LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs
by: Hou, Zhinan, et al.
Published: (2026)
by: Hou, Zhinan, et al.
Published: (2026)
Optimal Program Synthesis via Abstract Interpretation
by: Mell, Stephen, et al.
Published: (2026)
by: Mell, Stephen, et al.
Published: (2026)
Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning
by: Ding, Jianglin, et al.
Published: (2025)
by: Ding, Jianglin, et al.
Published: (2025)
Teaching LLMs to Think Mathematically: A Critical Study of Decision-Making via Optimization
by: Abdel-Rahman, Mohammad J., et al.
Published: (2025)
by: Abdel-Rahman, Mohammad J., et al.
Published: (2025)
Safety-Critical Control with Guaranteed Lipschitz Continuity via Filtered Control Barrier Functions
by: Liu, Shuo, et al.
Published: (2025)
by: Liu, Shuo, et al.
Published: (2025)
Cognitive Training for Language Models: Towards General Capabilities via Cross-Entropy Games
by: Hongler, Clément, et al.
Published: (2026)
by: Hongler, Clément, et al.
Published: (2026)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
A Mathematics-Inspired Learning-to-Optimize Framework for Decentralized Optimization
by: He, Yutong, et al.
Published: (2024)
by: He, Yutong, et al.
Published: (2024)
Stochastic Approximation with Delayed Updates: Finite-Time Rates under Markovian Sampling
by: Adibi, Arman, et al.
Published: (2024)
by: Adibi, Arman, et al.
Published: (2024)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
by: Anupam, Sagnik, et al.
Published: (2025)
by: Anupam, Sagnik, et al.
Published: (2025)
Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
by: Skifstad, Julian, et al.
Published: (2026)
by: Skifstad, Julian, et al.
Published: (2026)
Scaled Relative Graph of Normal Matrices
by: Huang, Xinmeng, et al.
Published: (2019)
by: Huang, Xinmeng, et al.
Published: (2019)
Similar Items
-
Uncertainty in Language Models: Assessment through Rank-Calibration
by: Huang, Xinmeng, et al.
Published: (2024) -
Alignment of large language models with constrained learning
by: Zhang, Botong, et al.
Published: (2025) -
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024) -
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
by: Mitra, Aritra, et al.
Published: (2023) -
Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
by: Huang, Xinmeng, et al.
Published: (2023)