SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Chegini, Atoosa, Kazemi, Hamid, Mirzadeh, Iman, Yin, Dong, Horton, Maxwell, Nabi, Moin, Farajtabar, Mehrdad, Alizadeh, Keivan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025)
by: Shojaee, Parshin, et al.
Published: (2025)
Computational Bottlenecks of Training Small-scale Large Language Models
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
by: Alizadeh, Keivan, et al.
Published: (2024)
by: Alizadeh, Keivan, et al.
Published: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024)
by: Mirzadeh, Iman, et al.
Published: (2024)
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
by: Kazemi, Hamid, et al.
Published: (2026)
by: Kazemi, Hamid, et al.
Published: (2026)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
by: Joudaki, Amir, et al.
Published: (2025)
by: Joudaki, Amir, et al.
Published: (2025)
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
by: Chegini, Atoosa, et al.
Published: (2025)
by: Chegini, Atoosa, et al.
Published: (2025)
RePanda: Pandas-powered Tabular Verification and Reasoning
by: Chegini, Atoosa Malemir, et al.
Published: (2025)
by: Chegini, Atoosa Malemir, et al.
Published: (2025)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
by: Alizadeh, Keivan, et al.
Published: (2023)
by: Alizadeh, Keivan, et al.
Published: (2023)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
by: Alizadeh, Keivan, et al.
Published: (2026)
by: Alizadeh, Keivan, et al.
Published: (2026)
What do we learn from inverting CLIP models?
by: Kazemi, Hamid, et al.
Published: (2024)
by: Kazemi, Hamid, et al.
Published: (2024)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
by: Chegini, Atoosa, et al.
Published: (2026)
by: Chegini, Atoosa, et al.
Published: (2026)
Learning Private Representations through Entropy-based Adversarial Training
by: Klein, Tassilo, et al.
Published: (2025)
by: Klein, Tassilo, et al.
Published: (2025)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
by: Nishu, Kumari, et al.
Published: (2025)
by: Nishu, Kumari, et al.
Published: (2025)
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
by: Klein, Tassilo, et al.
Published: (2024)
by: Klein, Tassilo, et al.
Published: (2024)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
by: Li, Jeffrey, et al.
Published: (2025)
by: Li, Jeffrey, et al.
Published: (2025)
Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
by: Saberi, Mehrdad, et al.
Published: (2023)
by: Saberi, Mehrdad, et al.
Published: (2023)
Offline Constrained RLHF with Multiple Preference Oracles
by: Latham, Brenden, et al.
Published: (2026)
by: Latham, Brenden, et al.
Published: (2026)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023)
by: Lin, Yong, et al.
Published: (2023)
No learning rates needed: Introducing SALSA -- Stable Armijo Line Search Adaptation
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
Improving Robustness of Foundation Models in Domain Adaptation with Soup-Adapters
by: Roschkowski, Marco
Published: (2025)
by: Roschkowski, Marco
Published: (2025)
Spectral Souping: A Unified Framework for Online Preference Alignment
by: Chow, Yinlam, et al.
Published: (2026)
by: Chow, Yinlam, et al.
Published: (2026)
OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
by: Huang, Chen, et al.
Published: (2025)
by: Huang, Chen, et al.
Published: (2025)
SALSA-RL: Stability Analysis in the Latent Space of Actions for Reinforcement Learning
by: Li, Xuyang, et al.
Published: (2025)
by: Li, Xuyang, et al.
Published: (2025)
Explaining and Preventing Alignment Collapse in Iterative RLHF
by: Gauthier, Etienne, et al.
Published: (2026)
by: Gauthier, Etienne, et al.
Published: (2026)
Beyond Model Interpretability: Socio-Structural Explanations in Machine Learning
by: Smart, Andrew, et al.
Published: (2024)
by: Smart, Andrew, et al.
Published: (2024)
Revisiting the Past: Data Unlearning with Model State History
by: Rezaei, Keivan, et al.
Published: (2025)
by: Rezaei, Keivan, et al.
Published: (2025)
Solving the Inverse Alignment Problem for Efficient RLHF
by: Krishna, Shambhavi, et al.
Published: (2024)
by: Krishna, Shambhavi, et al.
Published: (2024)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
Beyond RLHF: A Unified Theoretical Framework of Alignment
by: Yun, Jihun, et al.
Published: (2025)
by: Yun, Jihun, et al.
Published: (2025)
Thinking into the Future: Latent Lookahead Training for Transformers
by: Noci, Lorenzo, et al.
Published: (2026)
by: Noci, Lorenzo, et al.
Published: (2026)
KV Prediction for Improved Time to First Token
by: Horton, Maxwell, et al.
Published: (2024)
by: Horton, Maxwell, et al.
Published: (2024)
SALSA: Single-pass Autoregressive LLM Structured Classification
by: Berdichevsky, Ruslan, et al.
Published: (2025)
by: Berdichevsky, Ruslan, et al.
Published: (2025)
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
by: Yang, Zhiqin, et al.
Published: (2026)
by: Yang, Zhiqin, et al.
Published: (2026)
A Deep Bayesian Convolutional Spiking Neural Network-based CAD system with Uncertainty Quantification for Medical Images Classification
by: Chegini, Mohaddeseh, et al.
Published: (2025)
by: Chegini, Mohaddeseh, et al.
Published: (2025)
Robust CNN-based Respiration Rate Estimation for Smartwatch PPG and IMU
by: Kazemi, Kianoosh, et al.
Published: (2024)
by: Kazemi, Kianoosh, et al.
Published: (2024)
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
Similar Items
-
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025) -
Computational Bottlenecks of Training Small-scale Large Language Models
by: Ashkboos, Saleh, et al.
Published: (2024) -
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024) -
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
by: Alizadeh, Keivan, et al.
Published: (2024) -
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024)