Efficient Refusal Ablation in LLM through Optimal Transport
Fuente:
arXiv
Saved in:
| Main Authors: | Nanfack, Geraldin, Belilovsky, Eugene, Dohmatob, Elvis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
by: Nanfack, Geraldin, et al.
Published: (2025)
by: Nanfack, Geraldin, et al.
Published: (2025)
From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation
by: Nanfack, Geraldin, et al.
Published: (2024)
by: Nanfack, Geraldin, et al.
Published: (2024)
Test Time Adaptation Using Adaptive Quantile Recalibration
by: Mehrbod, Paria, et al.
Published: (2025)
by: Mehrbod, Paria, et al.
Published: (2025)
Model Collapse Demystified: The Case of Regression
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
by: Feng, Yunzhen, et al.
Published: (2024)
by: Feng, Yunzhen, et al.
Published: (2024)
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024)
by: Bayat, Reza, et al.
Published: (2024)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
by: Nabli, Adel, et al.
Published: (2024)
by: Nabli, Adel, et al.
Published: (2024)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Scaling Laws for Associative Memories
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Accelerating Training with Neuron Interaction and Nowcasting Networks
by: Knyazev, Boris, et al.
Published: (2024)
by: Knyazev, Boris, et al.
Published: (2024)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
Optimal Transport for LLM Reward Modeling from Noisy Preference
by: Pan, Licheng, et al.
Published: (2026)
by: Pan, Licheng, et al.
Published: (2026)
Unsupervised Anomaly Detection through Mass Repulsing Optimal Transport
by: Montesuma, Eduardo Fernandes, et al.
Published: (2025)
by: Montesuma, Eduardo Fernandes, et al.
Published: (2025)
Optimal Transport for Domain Adaptation through Gaussian Mixture Models
by: Montesuma, Eduardo Fernandes, et al.
Published: (2024)
by: Montesuma, Eduardo Fernandes, et al.
Published: (2024)
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
by: Pasand, Ali Saheb, et al.
Published: (2025)
by: Pasand, Ali Saheb, et al.
Published: (2025)
Harmony in Diversity: Merging Neural Networks with Canonical Correlation Analysis
by: Horoi, Stefan, et al.
Published: (2024)
by: Horoi, Stefan, et al.
Published: (2024)
Low-Rank Optimal Transport through Factor Relaxation with Latent Coupling
by: Halmos, Peter, et al.
Published: (2024)
by: Halmos, Peter, et al.
Published: (2024)
Imitation Learning from Observation through Optimal Transport
by: Chang, Wei-Di, et al.
Published: (2023)
by: Chang, Wei-Di, et al.
Published: (2023)
BEACON: Bayesian Optimal Stopping for Efficient LLM Sampling
by: Wan, Guangya, et al.
Published: (2025)
by: Wan, Guangya, et al.
Published: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
by: Huang, Jingkai, et al.
Published: (2026)
by: Huang, Jingkai, et al.
Published: (2026)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
by: Miahi, Erfan, et al.
Published: (2026)
by: Miahi, Erfan, et al.
Published: (2026)
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
by: Rahimi, Eliron, et al.
Published: (2026)
by: Rahimi, Eliron, et al.
Published: (2026)
Distributional Counterfactual Explanations With Optimal Transport
by: You, Lei, et al.
Published: (2024)
by: You, Lei, et al.
Published: (2024)
Displacement-Sparse Neural Optimal Transport
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Ablation Based Counterfactuals
by: Dai, Zheng, et al.
Published: (2024)
by: Dai, Zheng, et al.
Published: (2024)
On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration
by: Petrov, Valentin
Published: (2026)
by: Petrov, Valentin
Published: (2026)
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers
by: Kasneci, Gjergji, et al.
Published: (2024)
by: Kasneci, Gjergji, et al.
Published: (2024)
Amortized Optimal Transport from Sliced Potentials
by: Truong, Minh-Phuc, et al.
Published: (2026)
by: Truong, Minh-Phuc, et al.
Published: (2026)
Is Optimal Transport Necessary for Inverse Reinforcement Learning?
by: Dong, Zixuan, et al.
Published: (2025)
by: Dong, Zixuan, et al.
Published: (2025)
Counterfactual Identifiability via Dynamic Optimal Transport
by: Ribeiro, Fabio De Sousa, et al.
Published: (2025)
by: Ribeiro, Fabio De Sousa, et al.
Published: (2025)
Similar Items
-
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024) -
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
by: Nanfack, Geraldin, et al.
Published: (2025) -
From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation
by: Nanfack, Geraldin, et al.
Published: (2024) -
Test Time Adaptation Using Adaptive Quantile Recalibration
by: Mehrbod, Paria, et al.
Published: (2025) -
Model Collapse Demystified: The Case of Regression
by: Dohmatob, Elvis, et al.
Published: (2024)