Does Refusal Training in LLMs Generalize to the Past Tense?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Andriushchenko, Maksym, Flammarion, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is In-Context Learning Sufficient for Instruction Following in LLMs?
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
von: Fan, Dongyang, et al.
Veröffentlicht: (2026)
von: Fan, Dongyang, et al.
Veröffentlicht: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Decomposing and Measuring Evaluation Awareness
von: Li, Changling, et al.
Veröffentlicht: (2026)
von: Li, Changling, et al.
Veröffentlicht: (2026)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
von: Goel, Shashwat, et al.
Veröffentlicht: (2026)
von: Goel, Shashwat, et al.
Veröffentlicht: (2026)
QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
von: Qin, Jeremy, et al.
Veröffentlicht: (2026)
von: Qin, Jeremy, et al.
Veröffentlicht: (2026)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Where Do Reasoning Models Refuse?
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
von: Freeman, Joshua, et al.
Veröffentlicht: (2024)
von: Freeman, Joshua, et al.
Veröffentlicht: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
von: Mamun, Md Abdullah Al, et al.
Veröffentlicht: (2025)
von: Mamun, Md Abdullah Al, et al.
Veröffentlicht: (2025)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Does Biomedical Training Lead to Better Medical Performance?
von: Dada, Amin, et al.
Veröffentlicht: (2024)
von: Dada, Amin, et al.
Veröffentlicht: (2024)
Deep Language Geometry: Constructing a Metric Space from LLM Weights
von: Shamrai, Maksym, et al.
Veröffentlicht: (2025)
von: Shamrai, Maksym, et al.
Veröffentlicht: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
von: Frank, Gregory N.
Veröffentlicht: (2026)
von: Frank, Gregory N.
Veröffentlicht: (2026)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
Learning Algorithms in the Limit
von: Papazov, Hristo, et al.
Veröffentlicht: (2025)
von: Papazov, Hristo, et al.
Veröffentlicht: (2025)
Does Training on Synthetic Data Make Models Less Robust?
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
How to Train Data-Efficient LLMs
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
von: Zhao, Hao, et al.
Veröffentlicht: (2024)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
Continuous Approximations for Improving Quantization Aware Training of LLMs
von: Li, He, et al.
Veröffentlicht: (2024)
von: Li, He, et al.
Veröffentlicht: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
Thoth: Mid-Training Bridges LLMs to Time Series Understanding
von: Lin, Jiafeng, et al.
Veröffentlicht: (2026)
von: Lin, Jiafeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Is In-Context Learning Sufficient for Instruction Following in LLMs?
von: Zhao, Hao, et al.
Veröffentlicht: (2024) -
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024) -
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
von: Rando, Javier, et al.
Veröffentlicht: (2024) -
HalluHard: A Hard Multi-Turn Hallucination Benchmark
von: Fan, Dongyang, et al.
Veröffentlicht: (2026) -
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)