Differentially Private Steering for Large Language Model Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Goel, Anmol, Hu, Yaxi, Gurevych, Iryna, Sanyal, Amartya |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Auditing Language Model Unlearning via Information Decomposition
par: Goel, Anmol, et autres
Publié: (2026)
par: Goel, Anmol, et autres
Publié: (2026)
Provable Privacy with Non-Private Pre-Processing
par: Hu, Yaxi, et autres
Publié: (2024)
par: Hu, Yaxi, et autres
Publié: (2024)
How Quantization Shapes Bias in Large Language Models
par: Marcuzzi, Federico, et autres
Publié: (2025)
par: Marcuzzi, Federico, et autres
Publié: (2025)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
par: Tamoyan, Hovhannes, et autres
Publié: (2025)
par: Tamoyan, Hovhannes, et autres
Publié: (2025)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
par: Rizvi, Md Imbesat Hassan, et autres
Publié: (2024)
par: Rizvi, Md Imbesat Hassan, et autres
Publié: (2024)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
par: Geng, Jiahui, et autres
Publié: (2025)
par: Geng, Jiahui, et autres
Publié: (2025)
Online Learning and Unlearning
par: Hu, Yaxi, et autres
Publié: (2025)
par: Hu, Yaxi, et autres
Publié: (2025)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
par: Daheim, Nico, et autres
Publié: (2024)
par: Daheim, Nico, et autres
Publié: (2024)
Socratic Reasoning Improves Positive Text Rewriting
par: Goel, Anmol, et autres
Publié: (2024)
par: Goel, Anmol, et autres
Publié: (2024)
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
par: Rohweder, Jonas, et autres
Publié: (2026)
par: Rohweder, Jonas, et autres
Publié: (2026)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
par: Rizvi, Md Imbesat Hassan, et autres
Publié: (2025)
par: Rizvi, Md Imbesat Hassan, et autres
Publié: (2025)
LoRA and Privacy: When Random Projections Help (and When They Don't)
par: Hu, Yaxi, et autres
Publié: (2026)
par: Hu, Yaxi, et autres
Publié: (2026)
Differentially Private Zeroth-Order Methods for Scalable Large Language Model Finetuning
par: Liu, Z, et autres
Publié: (2024)
par: Liu, Z, et autres
Publié: (2024)
An Iterative Algorithm for Differentially Private $k$-PCA with Adaptive Noise
par: Düngler, Johanna, et autres
Publié: (2025)
par: Düngler, Johanna, et autres
Publié: (2025)
Differentially Private Next-Token Prediction of Large Language Models
par: Flemings, James, et autres
Publié: (2024)
par: Flemings, James, et autres
Publié: (2024)
Compositional Steering of Large Language Models with Steering Tokens
par: Radevski, Gorjan, et autres
Publié: (2026)
par: Radevski, Gorjan, et autres
Publié: (2026)
Differentially Private Tabular Data Synthesis using Large Language Models
par: Tran, Toan V., et autres
Publié: (2024)
par: Tran, Toan V., et autres
Publié: (2024)
Uncertainty-Aware Decoding with Minimum Bayes Risk
par: Daheim, Nico, et autres
Publié: (2025)
par: Daheim, Nico, et autres
Publié: (2025)
ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
par: Basch, Serwar, et autres
Publié: (2025)
par: Basch, Serwar, et autres
Publié: (2025)
Towards Automated Error Discovery: A Study in Conversational AI
par: Petrak, Dominic, et autres
Publié: (2025)
par: Petrak, Dominic, et autres
Publié: (2025)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
par: Thareja, Rushil, et autres
Publié: (2025)
par: Thareja, Rushil, et autres
Publié: (2025)
On the Growth of Mistakes in Differentially Private Online Learning: A Lower Bound Perspective
par: Dmitriev, Daniil, et autres
Publié: (2024)
par: Dmitriev, Daniil, et autres
Publié: (2024)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
par: Cao, Yuanpu, et autres
Publié: (2024)
par: Cao, Yuanpu, et autres
Publié: (2024)
Model Merging by Uncertainty-Based Gradient Matching
par: Daheim, Nico, et autres
Publié: (2023)
par: Daheim, Nico, et autres
Publié: (2023)
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
par: Li, Chen, et autres
Publié: (2026)
par: Li, Chen, et autres
Publié: (2026)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
par: Goel, Naman
Publié: (2023)
par: Goel, Naman
Publié: (2023)
LMO-DP: Optimizing the Randomization Mechanism for Differentially Private Fine-Tuning (Large) Language Models
par: Yang, Qin, et autres
Publié: (2024)
par: Yang, Qin, et autres
Publié: (2024)
Steering Language Models with Weight Arithmetic
par: Fierro, Constanza, et autres
Publié: (2025)
par: Fierro, Constanza, et autres
Publié: (2025)
Steering Language Models With Activation Engineering
par: Turner, Alexander Matt, et autres
Publié: (2023)
par: Turner, Alexander Matt, et autres
Publié: (2023)
Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language Models
par: Mekala, Anmol, et autres
Publié: (2024)
par: Mekala, Anmol, et autres
Publié: (2024)
Attribute or Abstain: Large Language Models as Long Document Assistants
par: Buchmann, Jan, et autres
Publié: (2024)
par: Buchmann, Jan, et autres
Publié: (2024)
Steering Large Language Models for Machine Translation Personalization
par: Scalena, Daniel, et autres
Publié: (2025)
par: Scalena, Daniel, et autres
Publié: (2025)
DP-OPD: Differentially Private On-Policy Distillation for Language Models
par: Khadem, Fatemeh, et autres
Publié: (2026)
par: Khadem, Fatemeh, et autres
Publié: (2026)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
par: Huang, Tiansheng, et autres
Publié: (2024)
par: Huang, Tiansheng, et autres
Publié: (2024)
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
par: Huo, Yifu, et autres
Publié: (2026)
par: Huo, Yifu, et autres
Publié: (2026)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
par: Sharma, Kartik, et autres
Publié: (2026)
par: Sharma, Kartik, et autres
Publié: (2026)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
par: Yang, Tianyu, et autres
Publié: (2024)
par: Yang, Tianyu, et autres
Publié: (2024)
CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
par: Huang, Weiyu, et autres
Publié: (2025)
par: Huang, Weiyu, et autres
Publié: (2025)
FedPDPO: Federated Personalized Direct Preference Optimization for Large Language Model Alignment
par: Zhu, Kewen, et autres
Publié: (2026)
par: Zhu, Kewen, et autres
Publié: (2026)
Discovering Implicit Large Language Model Alignment Objectives
par: Chen, Edward, et autres
Publié: (2026)
par: Chen, Edward, et autres
Publié: (2026)
Documents similaires
-
Auditing Language Model Unlearning via Information Decomposition
par: Goel, Anmol, et autres
Publié: (2026) -
Provable Privacy with Non-Private Pre-Processing
par: Hu, Yaxi, et autres
Publié: (2024) -
How Quantization Shapes Bias in Large Language Models
par: Marcuzzi, Federico, et autres
Publié: (2025) -
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
par: Tamoyan, Hovhannes, et autres
Publié: (2025) -
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
par: Rizvi, Md Imbesat Hassan, et autres
Publié: (2024)