Resurrecting saturated LLM benchmarks with adversarial encoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ivanov, Igor, Volkov, Dmitrii |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Badllama 3: removing safety finetuning from Llama 3 in minutes
von: Volkov, Dmitrii
Veröffentlicht: (2024)
von: Volkov, Dmitrii
Veröffentlicht: (2024)
BadGPT-4o: stripping safety finetuning from GPT models
von: Krupkina, Ekaterina, et al.
Veröffentlicht: (2024)
von: Krupkina, Ekaterina, et al.
Veröffentlicht: (2024)
The role of positional encodings in the ARC benchmark
von: Costa, Guilherme H. Bandeira, et al.
Veröffentlicht: (2025)
von: Costa, Guilherme H. Bandeira, et al.
Veröffentlicht: (2025)
Robust NAS under adversarial training: benchmark, theory, and beyond
von: Wu, Yongtao, et al.
Veröffentlicht: (2024)
von: Wu, Yongtao, et al.
Veröffentlicht: (2024)
The Resurrection of the ReLU
von: Horuz, Coşku Can, et al.
Veröffentlicht: (2025)
von: Horuz, Coşku Can, et al.
Veröffentlicht: (2025)
Fresh in memory: Training-order recency is linearly encoded in language model activations
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Resurrecting Label Propagation for Graphs with Heterophily and Label Noise
von: Cheng, Yao, et al.
Veröffentlicht: (2023)
von: Cheng, Yao, et al.
Veröffentlicht: (2023)
Robust LLM safeguarding via refusal feature adversarial training
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Test-time RL alignment exposes task familiarity artifacts in LLM benchmarks
von: Wang, Kun, et al.
Veröffentlicht: (2026)
von: Wang, Kun, et al.
Veröffentlicht: (2026)
Countering adversarial evasion in regression analysis
von: Benfield, David, et al.
Veröffentlicht: (2025)
von: Benfield, David, et al.
Veröffentlicht: (2025)
Provable tradeoffs in adversarially robust classification
von: Dobriban, Edgar, et al.
Veröffentlicht: (2020)
von: Dobriban, Edgar, et al.
Veröffentlicht: (2020)
On adversarial training and the 1 Nearest Neighbor classifier
von: Hagai, Amir, et al.
Veröffentlicht: (2024)
von: Hagai, Amir, et al.
Veröffentlicht: (2024)
Deployment-complete benchmarking
von: Mansouri, El Mustapha, et al.
Veröffentlicht: (2026)
von: Mansouri, El Mustapha, et al.
Veröffentlicht: (2026)
Lookahead identification in adversarial bandits: accuracy and memory bounds
von: Brukhim, Nataly, et al.
Veröffentlicht: (2026)
von: Brukhim, Nataly, et al.
Veröffentlicht: (2026)
On robust overfitting: adversarial training induced distribution matters
von: Tian, Runzhi, et al.
Veröffentlicht: (2023)
von: Tian, Runzhi, et al.
Veröffentlicht: (2023)
LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system
von: Li, Huanyu, et al.
Veröffentlicht: (2025)
von: Li, Huanyu, et al.
Veröffentlicht: (2025)
Can Go AIs be adversarially robust?
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
On damage of interpolation to adversarial robustness in regression
von: Peng, Jingfu, et al.
Veröffentlicht: (2026)
von: Peng, Jingfu, et al.
Veröffentlicht: (2026)
Featuremetric benchmarking: Quantum computer benchmarks based on circuit features
von: Proctor, Timothy, et al.
Veröffentlicht: (2025)
von: Proctor, Timothy, et al.
Veröffentlicht: (2025)
Online combinatorial optimization with stochastic decision sets and adversarial losses
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
von: Neu, Gergely, et al.
Veröffentlicht: (2026)
Best of both worlds: Stochastic & adversarial best-arm identification
von: Abbasi-Yadkori, Yasin, et al.
Veröffentlicht: (2026)
von: Abbasi-Yadkori, Yasin, et al.
Veröffentlicht: (2026)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
NODE-AdvGAN: Improving the transferability and perceptual similarity of adversarial examples by dynamic-system-driven adversarial generative model
von: Xie, Xinheng, et al.
Veröffentlicht: (2024)
von: Xie, Xinheng, et al.
Veröffentlicht: (2024)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
A unifying Bayesian framework for adversarial robustness
von: Arce, Pablo G., et al.
Veröffentlicht: (2025)
von: Arce, Pablo G., et al.
Veröffentlicht: (2025)
The role of class encoding in neural collapse
von: Massion, Bastien, et al.
Veröffentlicht: (2026)
von: Massion, Bastien, et al.
Veröffentlicht: (2026)
Does simple trump complex? Comparing strategies for adversarial robustness in DNNs
von: Brooks, William, et al.
Veröffentlicht: (2025)
von: Brooks, William, et al.
Veröffentlicht: (2025)
Masked adversarial neural network for cell type deconvolution in spatial transcriptomics
von: Huang, Lin, et al.
Veröffentlicht: (2024)
von: Huang, Lin, et al.
Veröffentlicht: (2024)
On the use of adversarial validation for quantifying dissimilarity in geospatial machine learning prediction
von: Wang, Yanwen, et al.
Veröffentlicht: (2024)
von: Wang, Yanwen, et al.
Veröffentlicht: (2024)
Generalization ability and Vulnerabilities to adversarial perturbations: Two sides of the same coin
von: Lee, Jung Hoon, et al.
Veröffentlicht: (2022)
von: Lee, Jung Hoon, et al.
Veröffentlicht: (2022)
Revoking Amnesia: RL-based Trajectory Optimization to Resurrect Erased Concepts in Diffusion Models
von: Gao, Daiheng, et al.
Veröffentlicht: (2025)
von: Gao, Daiheng, et al.
Veröffentlicht: (2025)
Concept activation vectors: a unifying view and adversarial attacks
von: Schnoor, Ekkehard, et al.
Veröffentlicht: (2025)
von: Schnoor, Ekkehard, et al.
Veröffentlicht: (2025)
Revealing data leakage in protein interaction benchmarks
von: Bushuiev, Anton, et al.
Veröffentlicht: (2024)
von: Bushuiev, Anton, et al.
Veröffentlicht: (2024)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
von: Reworr, et al.
Veröffentlicht: (2024)
von: Reworr, et al.
Veröffentlicht: (2024)
SHLIME: Foiling adversarial attacks fooling SHAP and LIME
von: Chauhan, Sam, et al.
Veröffentlicht: (2025)
von: Chauhan, Sam, et al.
Veröffentlicht: (2025)
A look at adversarial attacks on radio waveforms from discrete latent space
von: Garuso, Attanasia, et al.
Veröffentlicht: (2025)
von: Garuso, Attanasia, et al.
Veröffentlicht: (2025)
Deep MMD Gradient Flow without adversarial training
von: Galashov, Alexandre, et al.
Veröffentlicht: (2024)
von: Galashov, Alexandre, et al.
Veröffentlicht: (2024)
Robust estimation with Lasso when outputs are adversarially contaminated
von: Sasai, Takeyuki, et al.
Veröffentlicht: (2020)
von: Sasai, Takeyuki, et al.
Veröffentlicht: (2020)
Rates of convergence for density estimation with generative adversarial networks
von: Puchkin, Nikita, et al.
Veröffentlicht: (2021)
von: Puchkin, Nikita, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
Badllama 3: removing safety finetuning from Llama 3 in minutes
von: Volkov, Dmitrii
Veröffentlicht: (2024) -
BadGPT-4o: stripping safety finetuning from GPT models
von: Krupkina, Ekaterina, et al.
Veröffentlicht: (2024) -
The role of positional encodings in the ARC benchmark
von: Costa, Guilherme H. Bandeira, et al.
Veröffentlicht: (2025) -
Robust NAS under adversarial training: benchmark, theory, and beyond
von: Wu, Yongtao, et al.
Veröffentlicht: (2024) -
The Resurrection of the ReLU
von: Horuz, Coşku Can, et al.
Veröffentlicht: (2025)