Can Go AIs be adversarially robust?
Fuente:
arXiv
Guardado en:
| Autores principales: | Tseng, Tom, McLean, Euan, Pelrine, Kellin, Wang, Tony T., Gleave, Adam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploiting Novel GPT-4 APIs
por: Pelrine, Kellin, et al.
Publicado: (2023)
por: Pelrine, Kellin, et al.
Publicado: (2023)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
por: Struppek, Lukas, et al.
Publicado: (2026)
por: Struppek, Lukas, et al.
Publicado: (2026)
Scaling Trends for Data Poisoning in LLMs
por: Bowen, Dillon, et al.
Publicado: (2024)
por: Bowen, Dillon, et al.
Publicado: (2024)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
por: Kowal, Matthew, et al.
Publicado: (2026)
por: Kowal, Matthew, et al.
Publicado: (2026)
Language models are better than humans at next-token prediction
por: Shlegeris, Buck, et al.
Publicado: (2022)
por: Shlegeris, Buck, et al.
Publicado: (2022)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
por: Murphy, Brendan, et al.
Publicado: (2025)
por: Murphy, Brendan, et al.
Publicado: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
por: Cundy, Chris, et al.
Publicado: (2025)
por: Cundy, Chris, et al.
Publicado: (2025)
A Systematic Investigation of The RL-Jailbreaker in LLMs
por: Mohammedalamen, Montaser, et al.
Publicado: (2026)
por: Mohammedalamen, Montaser, et al.
Publicado: (2026)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
por: Taufeeque, Mohammad, et al.
Publicado: (2026)
por: Taufeeque, Mohammad, et al.
Publicado: (2026)
Enhancing robustness of data-driven SHM models: adversarial training with circle loss
por: Yang, Xiangli, et al.
Publicado: (2024)
por: Yang, Xiangli, et al.
Publicado: (2024)
RDI: An adversarial robustness evaluation metric for deep neural networks based on model statistical features
por: Song, Jialei, et al.
Publicado: (2025)
por: Song, Jialei, et al.
Publicado: (2025)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
por: Taufeeque, Mohammad, et al.
Publicado: (2025)
por: Taufeeque, Mohammad, et al.
Publicado: (2025)
Multi-Task Reinforcement Learning Enables Parameter Scaling
por: McLean, Reginald, et al.
Publicado: (2025)
por: McLean, Reginald, et al.
Publicado: (2025)
Curvature Dynamic Black-box Attack: revisiting adversarial robustness via dynamic curvature estimation
por: Sun, Peiran
Publicado: (2025)
por: Sun, Peiran
Publicado: (2025)
Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
por: Black, James R. M., et al.
Publicado: (2025)
por: Black, James R. M., et al.
Publicado: (2025)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
por: Bowen, Dillon, et al.
Publicado: (2025)
por: Bowen, Dillon, et al.
Publicado: (2025)
Scaling Trends in Language Model Robustness
por: Howe, Nikolaus, et al.
Publicado: (2024)
por: Howe, Nikolaus, et al.
Publicado: (2024)
STARC: A General Framework For Quantifying Differences Between Reward Functions
por: Skalse, Joar, et al.
Publicado: (2023)
por: Skalse, Joar, et al.
Publicado: (2023)
MXNorm: Reusing MXFP block scales for efficient tensor normalisation
por: McLean, Callum, et al.
Publicado: (2026)
por: McLean, Callum, et al.
Publicado: (2026)
Blending adversarial training and representation-conditional purification via aggregation improves adversarial robustness
por: Ballarin, Emanuele, et al.
Publicado: (2023)
por: Ballarin, Emanuele, et al.
Publicado: (2023)
$\texttt{MiniMol}$: A Parameter-Efficient Foundation Model for Molecular Learning
por: Kläser, Kerstin, et al.
Publicado: (2024)
por: Kläser, Kerstin, et al.
Publicado: (2024)
FRAUD-RLA: A new reinforcement learning adversarial attack against credit card fraud detection
por: Lunghi, Daniele, et al.
Publicado: (2025)
por: Lunghi, Daniele, et al.
Publicado: (2025)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
por: Lang, Leon, et al.
Publicado: (2024)
por: Lang, Leon, et al.
Publicado: (2024)
It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
por: Kowal, Matthew, et al.
Publicado: (2025)
por: Kowal, Matthew, et al.
Publicado: (2025)
Planning in a recurrent neural network that plays Sokoban
por: Taufeeque, Mohammad, et al.
Publicado: (2024)
por: Taufeeque, Mohammad, et al.
Publicado: (2024)
Limited but consistent gains in adversarial robustness by co-training object recognition models with human EEG
por: Guo, Manshan, et al.
Publicado: (2024)
por: Guo, Manshan, et al.
Publicado: (2024)
NODE-AdvGAN: Improving the transferability and perceptual similarity of adversarial examples by dynamic-system-driven adversarial generative model
por: Xie, Xinheng, et al.
Publicado: (2024)
por: Xie, Xinheng, et al.
Publicado: (2024)
Large language models can effectively convince people to believe conspiracies
por: Costello, Thomas H., et al.
Publicado: (2026)
por: Costello, Thomas H., et al.
Publicado: (2026)
Evaluating the robustness of adversarial defenses in malware detection systems
por: Jafari, Mostafa, et al.
Publicado: (2025)
por: Jafari, Mostafa, et al.
Publicado: (2025)
Deep MMD Gradient Flow without adversarial training
por: Galashov, Alexandre, et al.
Publicado: (2024)
por: Galashov, Alexandre, et al.
Publicado: (2024)
Robust NAS under adversarial training: benchmark, theory, and beyond
por: Wu, Yongtao, et al.
Publicado: (2024)
por: Wu, Yongtao, et al.
Publicado: (2024)
Missing value imputation with adversarial random forests -- MissARF
por: Golchian, Pegah, et al.
Publicado: (2025)
por: Golchian, Pegah, et al.
Publicado: (2025)
Fixed-point graph convolutional networks against adversarial attacks
por: Khan, Shakib, et al.
Publicado: (2025)
por: Khan, Shakib, et al.
Publicado: (2025)
Meta-World+: An Improved, Standardized, RL Benchmark
por: McLean, Reginald, et al.
Publicado: (2025)
por: McLean, Reginald, et al.
Publicado: (2025)
Deep generative models as an adversarial attack strategy for tabular machine learning
por: Dyrmishi, Salijona, et al.
Publicado: (2024)
por: Dyrmishi, Salijona, et al.
Publicado: (2024)
Properties that allow or prohibit transferability of adversarial attacks among quantized networks
por: Shrestha, Abhishek, et al.
Publicado: (2024)
por: Shrestha, Abhishek, et al.
Publicado: (2024)
How adversarial attacks can disrupt seemingly stable accurate classifiers
por: Sutton, Oliver J., et al.
Publicado: (2023)
por: Sutton, Oliver J., et al.
Publicado: (2023)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
por: Rivera, Mauricio, et al.
Publicado: (2024)
por: Rivera, Mauricio, et al.
Publicado: (2024)
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
por: Mazeika, Mantas, et al.
Publicado: (2025)
por: Mazeika, Mantas, et al.
Publicado: (2025)
Ejemplares similares
-
Exploiting Novel GPT-4 APIs
por: Pelrine, Kellin, et al.
Publicado: (2023) -
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
por: Struppek, Lukas, et al.
Publicado: (2026) -
Scaling Trends for Data Poisoning in LLMs
por: Bowen, Dillon, et al.
Publicado: (2024) -
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
por: Kowal, Matthew, et al.
Publicado: (2026) -
Language models are better than humans at next-token prediction
por: Shlegeris, Buck, et al.
Publicado: (2022)