Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Panfilov, Alexander, Kortukov, Evgenii, Nikolić, Kristina, Bethge, Matthias, Lapuschkin, Sebastian, Samek, Wojciech, Prabhu, Ameya, Andriushchenko, Maksym, Geiping, Jonas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
por: Terekhov, Mikhail, et al.
Publicado: (2025)
por: Terekhov, Mikhail, et al.
Publicado: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
por: Panfilov, Alexander, et al.
Publicado: (2025)
por: Panfilov, Alexander, et al.
Publicado: (2025)
ASIDE: Architectural Separation of Instructions and Data in Language Models
por: Zverev, Egor, et al.
Publicado: (2025)
por: Zverev, Egor, et al.
Publicado: (2025)
Great Models Think Alike and this Undermines AI Oversight
por: Goel, Shashwat, et al.
Publicado: (2025)
por: Goel, Shashwat, et al.
Publicado: (2025)
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
por: Sinha, Shiven, et al.
Publicado: (2025)
por: Sinha, Shiven, et al.
Publicado: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
por: Rank, Ben, et al.
Publicado: (2026)
por: Rank, Ben, et al.
Publicado: (2026)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
por: Skorobogat, Ronald, et al.
Publicado: (2026)
por: Skorobogat, Ronald, et al.
Publicado: (2026)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
por: Panfilov, Alexander, et al.
Publicado: (2026)
por: Panfilov, Alexander, et al.
Publicado: (2026)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
por: Pahde, Frederik, et al.
Publicado: (2025)
por: Pahde, Frederik, et al.
Publicado: (2025)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
por: Goel, Shashwat, et al.
Publicado: (2026)
por: Goel, Shashwat, et al.
Publicado: (2026)
Iterative Inference in a Chess-Playing Neural Network
por: Sandmann, Elias, et al.
Publicado: (2025)
por: Sandmann, Elias, et al.
Publicado: (2025)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
por: Erogullari, Eren, et al.
Publicado: (2025)
por: Erogullari, Eren, et al.
Publicado: (2025)
Explaining Predictive Uncertainty by Exposing Second-Order Effects
por: Bley, Florian, et al.
Publicado: (2024)
por: Bley, Florian, et al.
Publicado: (2024)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
por: Puri, Bruno, et al.
Publicado: (2025)
por: Puri, Bruno, et al.
Publicado: (2025)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
por: Dreyer, Maximilian, et al.
Publicado: (2023)
por: Dreyer, Maximilian, et al.
Publicado: (2023)
Mapping Post-Training Forgetting in Language Models at Scale
por: Harmon, Jackson, et al.
Publicado: (2025)
por: Harmon, Jackson, et al.
Publicado: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
por: Freeman, Joshua, et al.
Publicado: (2024)
por: Freeman, Joshua, et al.
Publicado: (2024)
From Weights to Activations: Is Steering the Next Frontier of Adaptation?
por: Ostermann, Simon, et al.
Publicado: (2026)
por: Ostermann, Simon, et al.
Publicado: (2026)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
por: Sinha, Shiven, et al.
Publicado: (2024)
por: Sinha, Shiven, et al.
Publicado: (2024)
CiteME: Can Language Models Accurately Cite Scientific Claims?
por: Press, Ori, et al.
Publicado: (2024)
por: Press, Ori, et al.
Publicado: (2024)
Human-Centered Evaluation of XAI Methods
por: Dawoud, Karam, et al.
Publicado: (2023)
por: Dawoud, Karam, et al.
Publicado: (2023)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
por: Boreiko, Valentyn, et al.
Publicado: (2024)
por: Boreiko, Valentyn, et al.
Publicado: (2024)
Are We Done with Object-Centric Learning?
por: Rubinstein, Alexander, et al.
Publicado: (2025)
por: Rubinstein, Alexander, et al.
Publicado: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
por: Chandak, Nikhil, et al.
Publicado: (2025)
por: Chandak, Nikhil, et al.
Publicado: (2025)
Scaling Open-Ended Reasoning to Predict the Future
por: Chandak, Nikhil, et al.
Publicado: (2025)
por: Chandak, Nikhil, et al.
Publicado: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
por: Dreyer, Maximilian, et al.
Publicado: (2024)
por: Dreyer, Maximilian, et al.
Publicado: (2024)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
por: Bouanani, Oussama, et al.
Publicado: (2026)
por: Bouanani, Oussama, et al.
Publicado: (2026)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
por: Becking, Daniel, et al.
Publicado: (2021)
por: Becking, Daniel, et al.
Publicado: (2021)
A Close Look at Decomposition-based XAI-Methods for Transformer Language Models
por: Arras, Leila, et al.
Publicado: (2025)
por: Arras, Leila, et al.
Publicado: (2025)
From Attribution to Action: A Human-Centered Application of Activation Steering
por: Labarta, Tobias, et al.
Publicado: (2026)
por: Labarta, Tobias, et al.
Publicado: (2026)
Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
por: Gururaj, Shreyas, et al.
Publicado: (2025)
por: Gururaj, Shreyas, et al.
Publicado: (2025)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
por: Bareeva, Dilyara, et al.
Publicado: (2024)
por: Bareeva, Dilyara, et al.
Publicado: (2024)
Building Trust in PINNs: Error Estimation through Finite Difference Methods
por: Krasowski, Aleksander, et al.
Publicado: (2026)
por: Krasowski, Aleksander, et al.
Publicado: (2026)
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
por: Ghosh, Adhiraj, et al.
Publicado: (2024)
por: Ghosh, Adhiraj, et al.
Publicado: (2024)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
por: Wiedemer, Thaddäus, et al.
Publicado: (2025)
por: Wiedemer, Thaddäus, et al.
Publicado: (2025)
Structural Compactness as a Complementary Criterion for Explanation Quality
por: Mesgari, Mohammad Mahdi, et al.
Publicado: (2026)
por: Mesgari, Mohammad Mahdi, et al.
Publicado: (2026)
Sparse, Efficient and Explainable Data Attribution with DualXDA
por: Yolcu, Galip Ümit, et al.
Publicado: (2024)
por: Yolcu, Galip Ümit, et al.
Publicado: (2024)
Ejemplares similares
-
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
por: Terekhov, Mikhail, et al.
Publicado: (2025) -
Capability-Based Scaling Trends for LLM-Based Red-Teaming
por: Panfilov, Alexander, et al.
Publicado: (2025) -
ASIDE: Architectural Separation of Instructions and Data in Language Models
por: Zverev, Egor, et al.
Publicado: (2025) -
Great Models Think Alike and this Undermines AI Oversight
por: Goel, Shashwat, et al.
Publicado: (2025) -
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
por: Sinha, Shiven, et al.
Publicado: (2025)