On the Failure of Topic-Matched Contrast Baselines in Multi-Directional Refusal Abliteration
Fuente:
arXiv
Guardado en:
| Autor principal: | Petrov, Valentin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
por: Shairah, Harethah Abu, et al.
Publicado: (2025)
por: Shairah, Harethah Abu, et al.
Publicado: (2025)
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
por: Piras, Giorgio, et al.
Publicado: (2025)
por: Piras, Giorgio, et al.
Publicado: (2025)
Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
por: Bhargav, Samaksh, et al.
Publicado: (2025)
por: Bhargav, Samaksh, et al.
Publicado: (2025)
Refusal in Language Models Is Mediated by a Single Direction
por: Arditi, Andy, et al.
Publicado: (2024)
por: Arditi, Andy, et al.
Publicado: (2024)
Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling
por: Kok-Shun, Brice Valentin, et al.
Publicado: (2026)
por: Kok-Shun, Brice Valentin, et al.
Publicado: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
por: Muhamed, Aashiq, et al.
Publicado: (2025)
por: Muhamed, Aashiq, et al.
Publicado: (2025)
Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
por: Grebe, Jonas Henry, et al.
Publicado: (2026)
por: Grebe, Jonas Henry, et al.
Publicado: (2026)
Efficient Refusal Ablation in LLM through Optimal Transport
por: Nanfack, Geraldin, et al.
Publicado: (2026)
por: Nanfack, Geraldin, et al.
Publicado: (2026)
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
por: Rahimi, Eliron, et al.
Publicado: (2026)
por: Rahimi, Eliron, et al.
Publicado: (2026)
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
por: Jia, Furong, et al.
Publicado: (2025)
por: Jia, Furong, et al.
Publicado: (2025)
Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
por: Zeng, Weili, et al.
Publicado: (2025)
por: Zeng, Weili, et al.
Publicado: (2025)
Where Do Reasoning Models Refuse?
por: Yamaguchi, Kureha, et al.
Publicado: (2025)
por: Yamaguchi, Kureha, et al.
Publicado: (2025)
Programming Refusal with Conditional Activation Steering
por: Lee, Bruce W., et al.
Publicado: (2024)
por: Lee, Bruce W., et al.
Publicado: (2024)
Integrating Distribution Matching into Semi-Supervised Contrastive Learning for Labeled and Unlabeled Data
por: Nakayama, Shogo, et al.
Publicado: (2026)
por: Nakayama, Shogo, et al.
Publicado: (2026)
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
por: Yang, Puning, et al.
Publicado: (2026)
por: Yang, Puning, et al.
Publicado: (2026)
CoSD: Collaborative Stance Detection with Contrastive Heterogeneous Topic Graph Learning
por: Cheng, Yinghan, et al.
Publicado: (2024)
por: Cheng, Yinghan, et al.
Publicado: (2024)
Model Failure or Data Corruption? Exploring Inconsistencies in Building Energy Ratings with Self-Supervised Contrastive Learning
por: Xiao, Qian, et al.
Publicado: (2024)
por: Xiao, Qian, et al.
Publicado: (2024)
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
por: Deng, Yonghong, et al.
Publicado: (2026)
por: Deng, Yonghong, et al.
Publicado: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
por: Hu, Xulin, et al.
Publicado: (2026)
por: Hu, Xulin, et al.
Publicado: (2026)
Merging Embedded Topics with Optimal Transport for Online Topic Modeling on Data Streams
por: Granese, Federica, et al.
Publicado: (2025)
por: Granese, Federica, et al.
Publicado: (2025)
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
por: Wu, Tongxi, et al.
Publicado: (2026)
por: Wu, Tongxi, et al.
Publicado: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
An Iterative Approach to Topic Modelling
por: Wong, Albert, et al.
Publicado: (2024)
por: Wong, Albert, et al.
Publicado: (2024)
Learning Multi-Agent Communication with Contrastive Learning
por: Lo, Yat Long, et al.
Publicado: (2023)
por: Lo, Yat Long, et al.
Publicado: (2023)
Simple Baselines are Competitive with Code Evolution
por: Gideoni, Yonatan, et al.
Publicado: (2026)
por: Gideoni, Yonatan, et al.
Publicado: (2026)
Integrated Influence: Data Attribution with Baseline
por: Yang, Linxiao, et al.
Publicado: (2025)
por: Yang, Linxiao, et al.
Publicado: (2025)
minimax: Efficient Baselines for Autocurricula in JAX
por: Jiang, Minqi, et al.
Publicado: (2023)
por: Jiang, Minqi, et al.
Publicado: (2023)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
por: Dabas, Mahavir, et al.
Publicado: (2025)
por: Dabas, Mahavir, et al.
Publicado: (2025)
DirectMultiStep: Direct Route Generation for Multistep Retrosynthesis
por: Shee, Yu, et al.
Publicado: (2024)
por: Shee, Yu, et al.
Publicado: (2024)
Tensor-Fused Multi-View Graph Contrastive Learning
por: Wu, Yujia, et al.
Publicado: (2024)
por: Wu, Yujia, et al.
Publicado: (2024)
Learnable Chernoff Baselines for Inference-Time Alignment
por: Madhow, Sunil, et al.
Publicado: (2026)
por: Madhow, Sunil, et al.
Publicado: (2026)
Topic Modeling and Link-Prediction for Material Property Discovery
por: Barron, Ryan C., et al.
Publicado: (2025)
por: Barron, Ryan C., et al.
Publicado: (2025)
Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
por: Zhang, Meiru, et al.
Publicado: (2026)
por: Zhang, Meiru, et al.
Publicado: (2026)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
por: Wollschläger, Tom, et al.
Publicado: (2025)
por: Wollschläger, Tom, et al.
Publicado: (2025)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
por: Sheng, Leheng, et al.
Publicado: (2025)
por: Sheng, Leheng, et al.
Publicado: (2025)
A Strong Baseline for Molecular Few-Shot Learning
por: Formont, Philippe, et al.
Publicado: (2024)
por: Formont, Philippe, et al.
Publicado: (2024)
TopicProphet: Prophesies on Temporal Topic Trends and Stocks
por: Kim, Olivia
Publicado: (2025)
por: Kim, Olivia
Publicado: (2025)
Refusing Safe Prompts for Multi-modal Large Language Models
por: Shao, Zedian, et al.
Publicado: (2024)
por: Shao, Zedian, et al.
Publicado: (2024)
Agentic Retrieval of Topics and Insights from Earnings Calls
por: Gupta, Anant, et al.
Publicado: (2025)
por: Gupta, Anant, et al.
Publicado: (2025)
Multi-agent Coordination via Flow Matching
por: Lee, Dongsu, et al.
Publicado: (2025)
por: Lee, Dongsu, et al.
Publicado: (2025)
Ejemplares similares
-
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
por: Shairah, Harethah Abu, et al.
Publicado: (2025) -
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
por: Piras, Giorgio, et al.
Publicado: (2025) -
Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
por: Bhargav, Samaksh, et al.
Publicado: (2025) -
Refusal in Language Models Is Mediated by a Single Direction
por: Arditi, Andy, et al.
Publicado: (2024) -
Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling
por: Kok-Shun, Brice Valentin, et al.
Publicado: (2026)