The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Fuente:
arXiv
Gespeichert in:
Ähnliche Einträge
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
Tamper-Resistant Safeguards for Open-Weight LLMs
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024)
von: Ren, Richard, et al.
Veröffentlicht: (2024)
Foundation models may exhibit staged progression in novel CBRN threat disclosure
von: Esvelt, Kevin M
Veröffentlicht: (2025)
von: Esvelt, Kevin M
Veröffentlicht: (2025)
FedSelect: Customized Selection of Parameters for Fine-Tuning during Personalized Federated Learning
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2023)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2023)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025)
von: Ren, Richard, et al.
Veröffentlicht: (2025)
Reducing Political Manipulation with Consistency Training
von: Phan, Long, et al.
Veröffentlicht: (2026)
von: Phan, Long, et al.
Veröffentlicht: (2026)
TextQuests: How Good are LLMs at Text-Based Video Games?
von: Phan, Long, et al.
Veröffentlicht: (2025)
von: Phan, Long, et al.
Veröffentlicht: (2025)
Thin Bridges for Drug Text Alignment: Lightweight Contrastive Learning for Target Specific Drug Retrieval
von: Tupakula, Mallikarjuna
Veröffentlicht: (2025)
von: Tupakula, Mallikarjuna
Veröffentlicht: (2025)
Aggressive Compression Enables LLM Weight Theft
von: Brown, Davis, et al.
Veröffentlicht: (2026)
von: Brown, Davis, et al.
Veröffentlicht: (2026)
FedSelect: Personalized Federated Learning with Customized Selection of Parameters for Fine-Tuning
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
COVID-19 Vaccines in the Pediatric Population: A Focus on Cardiac Patients
von: Ghena Lababidi, et al.
Veröffentlicht: (2024)
von: Ghena Lababidi, et al.
Veröffentlicht: (2024)
Reasons to Doubt the Impact of AI Risk Evaluations
von: Mukobi, Gabriel
Veröffentlicht: (2024)
von: Mukobi, Gabriel
Veröffentlicht: (2024)
Perception-Driven Bias Detection in Machine Learning via Crowdsourced Visual Judgment
von: Tupakula, Chirudeep, et al.
Veröffentlicht: (2025)
von: Tupakula, Chirudeep, et al.
Veröffentlicht: (2025)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
Bare Feet in the Ballroom: The First Demonstration in Australia of Dalcroze Eurhythmics, 1919
von: Oam, Joan Pope
Veröffentlicht: (2020)
von: Oam, Joan Pope
Veröffentlicht: (2020)
Kas yra politika?
von: Alvydas Jokubaitis
Veröffentlicht: (2022)
von: Alvydas Jokubaitis
Veröffentlicht: (2022)
Immanuelio Kanto iššūkis politikos mokslui
von: Alvydas Jokubaitis
Veröffentlicht: (2021)
von: Alvydas Jokubaitis
Veröffentlicht: (2021)
Politikos mokslo romantizmas
von: Alvydas Jokubaitis
Veröffentlicht: (2019)
von: Alvydas Jokubaitis
Veröffentlicht: (2019)
Mokslinės politinės filosofijos virtimas spekuliatyvia istorijos filosofija
von: Linas Jokubaitis
Veröffentlicht: (2020)
von: Linas Jokubaitis
Veröffentlicht: (2020)
Moralumo iššūkis Carlo Schmitto politiškumo sampratai
von: Alvydas Jokubaitis
Veröffentlicht: (2020)
von: Alvydas Jokubaitis
Veröffentlicht: (2020)
Immanuelis Kantas ir politikos tikrumo klausimas
von: Alvydas Jokubaitis
Veröffentlicht: (2021)
von: Alvydas Jokubaitis
Veröffentlicht: (2021)
Politinė Stasio Šalkauskio kultūros filosofijos prasmė
von: Alvydas Jokubaitis
Veröffentlicht: (2020)
von: Alvydas Jokubaitis
Veröffentlicht: (2020)
Jet-Density of Finite-Gap Solutions for Classes of BKM Systems
von: Quaschner, Manuel, et al.
Veröffentlicht: (2026)
von: Quaschner, Manuel, et al.
Veröffentlicht: (2026)
Conocimiento y método en Descartes, Pascal y Leibniz
von: Josep M. Basart Muñoz
Veröffentlicht: (2004)
von: Josep M. Basart Muñoz
Veröffentlicht: (2004)
Superintelligence Strategy: Expert Version
von: Hendrycks, Dan, et al.
Veröffentlicht: (2025)
von: Hendrycks, Dan, et al.
Veröffentlicht: (2025)
Private equity acquisitions and product market decisions: Evidence from trademarks
von: Moazzam Khoja
Veröffentlicht: (2025)
von: Moazzam Khoja
Veröffentlicht: (2025)
Complexity Classification of Product State Problems for Local Hamiltonians
von: Kallaugher, John, et al.
Veröffentlicht: (2024)
von: Kallaugher, John, et al.
Veröffentlicht: (2024)
The effect of confession evidence on conviction, and considering alternative scenarios as remedy in a sample of police officers
von: Neville Niccolson, et al.
Veröffentlicht: (2024)
von: Neville Niccolson, et al.
Veröffentlicht: (2024)
Data-Efficient Ensemble Weather Forecasting with Diffusion Models
von: Valencia, Kevin, et al.
Veröffentlicht: (2025)
von: Valencia, Kevin, et al.
Veröffentlicht: (2025)
Explaining Snowball-in-hell Phenomena in Heavy-ion Collisions Using a Novel Thermodynamic Variable
von: Braaten, Eric, et al.
Veröffentlicht: (2024)
von: Braaten, Eric, et al.
Veröffentlicht: (2024)
Video Text Preservation with Synthetic Text-Rich Videos
von: Liu, Ziyang, et al.
Veröffentlicht: (2025)
von: Liu, Ziyang, et al.
Veröffentlicht: (2025)
LOOM: Personalized Learning Informed by Daily LLM Conversations Toward Long-Term Mastery via a Dynamic Learner Memory Graph
von: Cui, Justin, et al.
Veröffentlicht: (2025)
von: Cui, Justin, et al.
Veröffentlicht: (2025)
Introduction to AI Safety, Ethics, and Society
von: Hendrycks, Dan
Veröffentlicht: (2024)
von: Hendrycks, Dan
Veröffentlicht: (2024)
Introduction to AI Safety, Ethics, and Society
von: Hendrycks, Dan
Veröffentlicht: (2024)
von: Hendrycks, Dan
Veröffentlicht: (2024)
Variationality of conformal geodesics in dimension 3
von: Kruglikov, Boris, et al.
Veröffentlicht: (2024)
von: Kruglikov, Boris, et al.
Veröffentlicht: (2024)
Leveraging Quantum Computing for Accelerated Classical Algorithms in Power Systems Optimization
von: Barrass, Rosemary, et al.
Veröffentlicht: (2025)
von: Barrass, Rosemary, et al.
Veröffentlicht: (2025)
On globally invariant Euler--Lagrange equations for curves
von: Kruglikov, Boris, et al.
Veröffentlicht: (2026)
von: Kruglikov, Boris, et al.
Veröffentlicht: (2026)
Measuring the Robustness of Reference-Free Dialogue Evaluation Systems
von: Vasselli, Justin, et al.
Veröffentlicht: (2025)
von: Vasselli, Justin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025) -
Tamper-Resistant Safeguards for Open-Weight LLMs
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024) -
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024) -
Foundation models may exhibit staged progression in novel CBRN threat disclosure
von: Esvelt, Kevin M
Veröffentlicht: (2025) -
FedSelect: Customized Selection of Parameters for Fine-Tuning during Personalized Federated Learning
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2023)