Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Black, James R. M., Hanke, Moritz S., Maiwald, Aaron, Hernandez-Boussard, Tina, Crook, Oliver M., Pannu, Jaspreet |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Securing Dual-Use Pathogen Data of Concern
by: Bloomfield, Doni, et al.
Published: (2026)
by: Bloomfield, Doni, et al.
Published: (2026)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
The effect of fine-tuning on language model toxicity
by: Hawkins, Will, et al.
Published: (2024)
by: Hawkins, Will, et al.
Published: (2024)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Deep literature reviews: an application of fine-tuned language models to migration research
by: Iacus, Stefano M., et al.
Published: (2025)
by: Iacus, Stefano M., et al.
Published: (2025)
Improving generative adversarial network inversion via fine-tuning GAN encoders
by: Yu, Cheng, et al.
Published: (2021)
by: Yu, Cheng, et al.
Published: (2021)
Retrieval and competition: how a protein foundation model starts a protein
by: Jedryszek, Piotr, et al.
Published: (2026)
by: Jedryszek, Piotr, et al.
Published: (2026)
Fundamentos de seguridad de redes / Eric Maiwald; trad. Efrén Alatorre Miguel
by: Maiwald, Eric
Published: (2005)
by: Maiwald, Eric
Published: (2005)
Out-of-distribution materials property prediction using adversarial learning based fine-tuning
by: Li, Qinyang, et al.
Published: (2024)
by: Li, Qinyang, et al.
Published: (2024)
How does fine-tuning improve sensorimotor representations in large language models?
by: Wu, Minghua, et al.
Published: (2026)
by: Wu, Minghua, et al.
Published: (2026)
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
by: Han, Yifu, et al.
Published: (2025)
by: Han, Yifu, et al.
Published: (2025)
The representation landscape of few-shot learning and fine-tuning in large language models
by: Doimo, Diego, et al.
Published: (2024)
by: Doimo, Diego, et al.
Published: (2024)
Stronger adversaries grow cheaper forests: online node-weighted Steiner problems
by: Borst, Sander, et al.
Published: (2024)
by: Borst, Sander, et al.
Published: (2024)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
FairEHR-CLP: Towards Fairness-Aware Clinical Predictions with Contrastive Learning in Multimodal Electronic Health Records
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
Can adversarial attacks by large language models be attributed?
by: Cebrian, Manuel, et al.
Published: (2024)
by: Cebrian, Manuel, et al.
Published: (2024)
RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
by: Zheng, Jingnan, et al.
Published: (2025)
by: Zheng, Jingnan, et al.
Published: (2025)
Stable and Steerable Sparse Autoencoders with Weight Regularization
by: Jedryszek, Piotr, et al.
Published: (2026)
by: Jedryszek, Piotr, et al.
Published: (2026)
Provable tradeoffs in adversarially robust classification
by: Dobriban, Edgar, et al.
Published: (2020)
by: Dobriban, Edgar, et al.
Published: (2020)
Can Go AIs be adversarially robust?
by: Tseng, Tom, et al.
Published: (2024)
by: Tseng, Tom, et al.
Published: (2024)
On damage of interpolation to adversarial robustness in regression
by: Peng, Jingfu, et al.
Published: (2026)
by: Peng, Jingfu, et al.
Published: (2026)
Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
by: Abdullah, Abdulhady Abas, et al.
Published: (2025)
by: Abdullah, Abdulhady Abas, et al.
Published: (2025)
Alternative weighting schemes for fine‐tuned extended similarity indices
by: Kenneth López Pérez, et al.
Published: (2024)
by: Kenneth López Pérez, et al.
Published: (2024)
MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Efficient and Private: Memorisation under differentially private parameter-efficient fine-tuning in language models
by: Ma, Olivia, et al.
Published: (2024)
by: Ma, Olivia, et al.
Published: (2024)
Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
by: Lyu, Y., et al.
Published: (2025)
by: Lyu, Y., et al.
Published: (2025)
On fine-tuning Boltz-2 for protein-protein affinity prediction
by: King, James, et al.
Published: (2025)
by: King, James, et al.
Published: (2025)
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
by: Maharjan, Jenish, et al.
Published: (2024)
by: Maharjan, Jenish, et al.
Published: (2024)
Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention Dynamics
by: Koubbi, Hugo, et al.
Published: (2024)
by: Koubbi, Hugo, et al.
Published: (2024)
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
by: Jung, Vincent, et al.
Published: (2024)
by: Jung, Vincent, et al.
Published: (2024)
Bayesian sample size determination using robust commensurate priors with interpretable discrepancy weights
by: Whitehead, Lou E., et al.
Published: (2024)
by: Whitehead, Lou E., et al.
Published: (2024)
Spectral regularization for adversarially-robust representation learning
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
On global robustness of an adversarial risk analysis solution
by: Jinming Yang, et al.
Published: (2024)
by: Jinming Yang, et al.
Published: (2024)
A unifying Bayesian framework for adversarial robustness
by: Arce, Pablo G., et al.
Published: (2025)
by: Arce, Pablo G., et al.
Published: (2025)
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
by: Xian, R. Patrick, et al.
Published: (2024)
by: Xian, R. Patrick, et al.
Published: (2024)
Blending adversarial training and representation-conditional purification via aggregation improves adversarial robustness
by: Ballarin, Emanuele, et al.
Published: (2023)
by: Ballarin, Emanuele, et al.
Published: (2023)
Learning Interpretable Point-Based Clinical Risk Scores via Direct Optimization
by: Cui, Ying, et al.
Published: (2026)
by: Cui, Ying, et al.
Published: (2026)
Limited but consistent gains in adversarial robustness by co-training object recognition models with human EEG
by: Guo, Manshan, et al.
Published: (2024)
by: Guo, Manshan, et al.
Published: (2024)
Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
by: Fuoli, Matteo, et al.
Published: (2025)
by: Fuoli, Matteo, et al.
Published: (2025)
Sequential sample size calculations and learning curves safeguard the robust development of a clinical prediction model for individuals
by: Legha, Amardeep, et al.
Published: (2025)
by: Legha, Amardeep, et al.
Published: (2025)
Similar Items
-
Securing Dual-Use Pathogen Data of Concern
by: Bloomfield, Doni, et al.
Published: (2026) -
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026) -
The effect of fine-tuning on language model toxicity
by: Hawkins, Will, et al.
Published: (2024) -
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024) -
Deep literature reviews: an application of fine-tuned language models to migration research
by: Iacus, Stefano M., et al.
Published: (2025)