Evaluating Defences against Unsafe Feedback in RLHF
Fuente:
arXiv
Salvato in:
| Autori principali: | Rosati, Domenic, Edkins, Giles, Raj, Harsh, Atanasov, David, Majumdar, Subhabrata, Rajendran, Janarthanan, Rudzicz, Frank, Sajjad, Hassan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Representation Noising: A Defence Mechanism Against Harmful Finetuning
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
Limits of Convergence-Rate Control for Open-Weight Safety
di: Rosati, Domenic, et al.
Pubblicazione: (2026)
di: Rosati, Domenic, et al.
Pubblicazione: (2026)
Semantic Consistency for Assuring Reliability of Large Language Models
di: Raj, Harsh, et al.
Pubblicazione: (2023)
di: Raj, Harsh, et al.
Pubblicazione: (2023)
Improving Consistency in Large Language Models through Chain of Guidance
di: Raj, Harsh, et al.
Pubblicazione: (2025)
di: Raj, Harsh, et al.
Pubblicazione: (2025)
Immunization against harmful fine-tuning attacks
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
Resolving Lexical Bias in Model Editing
di: Rizwan, Hammad, et al.
Pubblicazione: (2024)
di: Rizwan, Hammad, et al.
Pubblicazione: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
di: Lee, Harrison, et al.
Pubblicazione: (2023)
di: Lee, Harrison, et al.
Pubblicazione: (2023)
Consistency in Language Models: Current Landscape, Challenges, and Future Directions
di: Novikova, Jekaterina, et al.
Pubblicazione: (2025)
di: Novikova, Jekaterina, et al.
Pubblicazione: (2025)
LLM Library Learning Fails: A LEGO-Prover Case Study
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2025)
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2025)
Dependency Parsing is More Parameter-Efficient with Normalization
di: Gajo, Paolo, et al.
Pubblicazione: (2025)
di: Gajo, Paolo, et al.
Pubblicazione: (2025)
Long-form evaluation of model editing
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
LLMs Underperform Graph-Based Parsers on Supervised Relation Extraction for Complex Graphs
di: Gajo, Paolo, et al.
Pubblicazione: (2026)
di: Gajo, Paolo, et al.
Pubblicazione: (2026)
Localized LoRA: A Structured Low-Rank Approximation for Efficient Fine-Tuning
di: Barazandeh, Babak, et al.
Pubblicazione: (2025)
di: Barazandeh, Babak, et al.
Pubblicazione: (2025)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2024)
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2024)
Learning and Forgetting Unsafe Examples in Large Language Models
di: Zhao, Jiachen, et al.
Pubblicazione: (2023)
di: Zhao, Jiachen, et al.
Pubblicazione: (2023)
How to Evaluate Reward Models for RLHF
di: Frick, Evan, et al.
Pubblicazione: (2024)
di: Frick, Evan, et al.
Pubblicazione: (2024)
Interpreting the Effects of Quantization on LLMs
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
di: Singh, Manpreet, et al.
Pubblicazione: (2025)
Quantifying the Capabilities of LLMs across Scale and Precision
di: Badshah, Sher, et al.
Pubblicazione: (2024)
di: Badshah, Sher, et al.
Pubblicazione: (2024)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
di: Ajwani, Rohan Deepak, et al.
Pubblicazione: (2024)
di: Ajwani, Rohan Deepak, et al.
Pubblicazione: (2024)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
di: Chaudhari, Shreyas, et al.
Pubblicazione: (2024)
di: Chaudhari, Shreyas, et al.
Pubblicazione: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
di: Dong, Hanze, et al.
Pubblicazione: (2024)
di: Dong, Hanze, et al.
Pubblicazione: (2024)
Failure Modes of Maximum Entropy RLHF
di: Çağatan, Ömer Veysel, et al.
Pubblicazione: (2025)
di: Çağatan, Ömer Veysel, et al.
Pubblicazione: (2025)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
di: Hu, Jian, et al.
Pubblicazione: (2024)
di: Hu, Jian, et al.
Pubblicazione: (2024)
Mastering Memory Tasks with World Models
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
Intelligent Switching for Reset-Free RL
di: Patil, Darshan, et al.
Pubblicazione: (2024)
di: Patil, Darshan, et al.
Pubblicazione: (2024)
LLMGuard: Guarding Against Unsafe LLM Behavior
di: Goyal, Shubh, et al.
Pubblicazione: (2024)
di: Goyal, Shubh, et al.
Pubblicazione: (2024)
Solving the Inverse Alignment Problem for Efficient RLHF
di: Krishna, Shambhavi, et al.
Pubblicazione: (2024)
di: Krishna, Shambhavi, et al.
Pubblicazione: (2024)
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
di: Rohweder, Jonas, et al.
Pubblicazione: (2026)
di: Rohweder, Jonas, et al.
Pubblicazione: (2026)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
Group Robust Preference Optimization in Reward-free RLHF
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2024)
di: Ramesh, Shyam Sundhar, et al.
Pubblicazione: (2024)
Why Is RLHF Alignment Shallow? A Gradient Analysis
di: Young, Robin
Pubblicazione: (2026)
di: Young, Robin
Pubblicazione: (2026)
RLHF and IIA: Perverse Incentives
di: Xu, Wanqiao, et al.
Pubblicazione: (2023)
di: Xu, Wanqiao, et al.
Pubblicazione: (2023)
Reward-Robust RLHF in LLMs
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
di: Yu, Xuemin, et al.
Pubblicazione: (2026)
di: Yu, Xuemin, et al.
Pubblicazione: (2026)
Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
di: Si, Wai Man, et al.
Pubblicazione: (2026)
di: Si, Wai Man, et al.
Pubblicazione: (2026)
A Long Way to Go: Investigating Length Correlations in RLHF
di: Singhal, Prasann, et al.
Pubblicazione: (2023)
di: Singhal, Prasann, et al.
Pubblicazione: (2023)
Reward Model Overoptimisation in Iterated RLHF
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
Dataset Reset Policy Optimization for RLHF
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
Measuring memorization in RLHF for code completion
di: Pappu, Aneesh, et al.
Pubblicazione: (2024)
di: Pappu, Aneesh, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Representation Noising: A Defence Mechanism Against Harmful Finetuning
di: Rosati, Domenic, et al.
Pubblicazione: (2024) -
Limits of Convergence-Rate Control for Open-Weight Safety
di: Rosati, Domenic, et al.
Pubblicazione: (2026) -
Semantic Consistency for Assuring Reliability of Large Language Models
di: Raj, Harsh, et al.
Pubblicazione: (2023) -
Improving Consistency in Large Language Models through Chain of Guidance
di: Raj, Harsh, et al.
Pubblicazione: (2025) -
Immunization against harmful fine-tuning attacks
di: Rosati, Domenic, et al.
Pubblicazione: (2024)