PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Verma, Richa, Kulur, Bavish, Chawla, Sanjay, Ravindran, Balaraman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalized Adaptive Transfer Network: Enhancing Transfer Learning in Reinforcement Learning Across Domains
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics
by: Acharjee, Jashaswimalya, et al.
Published: (2026)
by: Acharjee, Jashaswimalya, et al.
Published: (2026)
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2025)
by: Burnwal, Returaj, et al.
Published: (2025)
OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2026)
by: Burnwal, Returaj, et al.
Published: (2026)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
by: Springer, Max, et al.
Published: (2026)
by: Springer, Max, et al.
Published: (2026)
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
by: Hossain, Saad, et al.
Published: (2025)
by: Hossain, Saad, et al.
Published: (2025)
Learning from Observation: A Survey of Recent Advances
by: Burnwal, Returaj, et al.
Published: (2025)
by: Burnwal, Returaj, et al.
Published: (2025)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
by: Qi, Xuan, et al.
Published: (2025)
by: Qi, Xuan, et al.
Published: (2025)
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
by: Taraghi, Mina, et al.
Published: (2025)
by: Taraghi, Mina, et al.
Published: (2025)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Alignment Dynamics in LLM Fine-Tuning
by: Huang, Yuhan, et al.
Published: (2026)
by: Huang, Yuhan, et al.
Published: (2026)
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
by: Parthasarathy, Ambreesh, et al.
Published: (2025)
Reward Sharpness-Aware Fine-Tuning for Diffusion Models
by: Kim, Kwanyoung, et al.
Published: (2026)
by: Kim, Kwanyoung, et al.
Published: (2026)
Safety Modulation: Enhancing Safety in Reinforcement Learning through Cost-Modulated Rewards
by: Zhang, Hanping, et al.
Published: (2025)
by: Zhang, Hanping, et al.
Published: (2025)
Hindsight PRIORs for Reward Learning from Human Preferences
by: Verma, Mudit, et al.
Published: (2024)
by: Verma, Mudit, et al.
Published: (2024)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
by: Afsharrad, Amirhossein, et al.
Published: (2026)
by: Afsharrad, Amirhossein, et al.
Published: (2026)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
by: Rajaram, Sara, et al.
Published: (2025)
by: Rajaram, Sara, et al.
Published: (2025)
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
by: Ravindran, Santhosh Kumar
Published: (2025)
by: Ravindran, Santhosh Kumar
Published: (2025)
Understanding and Preserving Safety in Fine-Tuned LLMs
by: Zhang, Jiawen, et al.
Published: (2026)
by: Zhang, Jiawen, et al.
Published: (2026)
FLAME: Towards Federated Fine-Tuning Large Language Models Through Adaptive SMoE
by: Le, Khiem, et al.
Published: (2025)
by: Le, Khiem, et al.
Published: (2025)
Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Implicit Federated In-context Learning For Task-Specific LLM Fine-Tuning
by: Li, Dongcheng, et al.
Published: (2025)
by: Li, Dongcheng, et al.
Published: (2025)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
by: Chen, Haoxian, et al.
Published: (2024)
by: Chen, Haoxian, et al.
Published: (2024)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
Representation Without Reward: A JEPA Audit for LLM Fine-Tuning
by: Sengupta, Biswa
Published: (2026)
by: Sengupta, Biswa
Published: (2026)
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
by: Liu, Zehua, et al.
Published: (2026)
by: Liu, Zehua, et al.
Published: (2026)
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
by: Lin, Baijiong, et al.
Published: (2025)
by: Lin, Baijiong, et al.
Published: (2025)
Fine-Tuning Language Models with Reward Learning on Policy
by: Lang, Hao, et al.
Published: (2024)
by: Lang, Hao, et al.
Published: (2024)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
by: Tejaswi, Atula, et al.
Published: (2026)
by: Tejaswi, Atula, et al.
Published: (2026)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
by: Bhambri, Siddhant, et al.
Published: (2023)
by: Bhambri, Siddhant, et al.
Published: (2023)
Course-Correction: Safety Alignment Using Synthetic Preferences
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Similar Items
-
Generalized Adaptive Transfer Network: Enhancing Transfer Learning in Reinforcement Learning Across Domains
by: Verma, Abhishek, et al.
Published: (2025) -
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
by: Verma, Abhishek, et al.
Published: (2025) -
Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics
by: Acharjee, Jashaswimalya, et al.
Published: (2026) -
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2025) -
OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2026)