Estimating Worst-Case Frontier Risks of Open-Weight LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Wallace, Eric, Watkins, Olivia, Wang, Miles, Chen, Kai, Koch, Chris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Differentially Private Worst-group Risk Minimization
por: Zhou, Xinyu, et al.
Publicado: (2024)
por: Zhou, Xinyu, et al.
Publicado: (2024)
Tamper-Resistant Safeguards for Open-Weight LLMs
por: Tamirisa, Rishub, et al.
Publicado: (2024)
por: Tamirisa, Rishub, et al.
Publicado: (2024)
How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness
por: Rossolini, Giulio
Publicado: (2026)
por: Rossolini, Giulio
Publicado: (2026)
Open Problems in Frontier AI Risk Management
por: Ziosi, Marta, et al.
Publicado: (2026)
por: Ziosi, Marta, et al.
Publicado: (2026)
Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
por: Aryal, Manish, et al.
Publicado: (2026)
por: Aryal, Manish, et al.
Publicado: (2026)
Risk Profiling and Modulation for LLMs
por: Wang, Yikai, et al.
Publicado: (2025)
por: Wang, Yikai, et al.
Publicado: (2025)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
por: Liang, Zhiyuan, et al.
Publicado: (2025)
por: Liang, Zhiyuan, et al.
Publicado: (2025)
NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
por: Fiotto-Kaufman, Jaden, et al.
Publicado: (2024)
por: Fiotto-Kaufman, Jaden, et al.
Publicado: (2024)
Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret
por: Zhong, Han, et al.
Publicado: (2023)
por: Zhong, Han, et al.
Publicado: (2023)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
por: O'Brien, Kyle, et al.
Publicado: (2025)
por: O'Brien, Kyle, et al.
Publicado: (2025)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
por: Guo, Chuan, et al.
Publicado: (2026)
por: Guo, Chuan, et al.
Publicado: (2026)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
por: Freeman, Joshua, et al.
Publicado: (2024)
por: Freeman, Joshua, et al.
Publicado: (2024)
Adversarial Training for Robust Coverage Network under Worst-case Facility Losses
por: Miao, Changhao, et al.
Publicado: (2026)
por: Miao, Changhao, et al.
Publicado: (2026)
NVLM: Open Frontier-Class Multimodal LLMs
por: Dai, Wenliang, et al.
Publicado: (2024)
por: Dai, Wenliang, et al.
Publicado: (2024)
Worst-Case Regret Bounds for Exploration via Randomized Value Functions
por: Russo, Daniel
Publicado: (2019)
por: Russo, Daniel
Publicado: (2019)
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
por: Lin, Yong, et al.
Publicado: (2025)
por: Lin, Yong, et al.
Publicado: (2025)
Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
por: Ke, Luoma, et al.
Publicado: (2024)
por: Ke, Luoma, et al.
Publicado: (2024)
OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data
por: Renda, Alana, et al.
Publicado: (2025)
por: Renda, Alana, et al.
Publicado: (2025)
The Path to Open Innovation: Peer-Review Under Fire (PRUF)
por: Billions, Ava, et al.
Publicado: (2025)
por: Billions, Ava, et al.
Publicado: (2025)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
por: Dickson, Craig
Publicado: (2025)
por: Dickson, Craig
Publicado: (2025)
Bounding the Worst-class Error: A Boosting Approach
por: Saito, Yuya, et al.
Publicado: (2023)
por: Saito, Yuya, et al.
Publicado: (2023)
Worst-case low-rank approximations
por: Fries, Anya, et al.
Publicado: (2026)
por: Fries, Anya, et al.
Publicado: (2026)
AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs
por: Hogan, Brendan R., et al.
Publicado: (2026)
por: Hogan, Brendan R., et al.
Publicado: (2026)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
por: Yang, Zeyu, et al.
Publicado: (2025)
por: Yang, Zeyu, et al.
Publicado: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024)
por: Luo, Yingsong, et al.
Publicado: (2024)
Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness
por: Park, Subeen, et al.
Publicado: (2025)
por: Park, Subeen, et al.
Publicado: (2025)
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
por: Wang, Miles, et al.
Publicado: (2026)
por: Wang, Miles, et al.
Publicado: (2026)
Foundations and Frontiers of Graph Learning Theory
por: Huang, Yu, et al.
Publicado: (2024)
por: Huang, Yu, et al.
Publicado: (2024)
Position: Weight Space Should Be a First-Class Generative AI Modality
por: Wang, Zhangyang, et al.
Publicado: (2026)
por: Wang, Zhangyang, et al.
Publicado: (2026)
Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost
por: Dennis, Simon, et al.
Publicado: (2026)
por: Dennis, Simon, et al.
Publicado: (2026)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
por: Zheng, Kunhao, et al.
Publicado: (2026)
por: Zheng, Kunhao, et al.
Publicado: (2026)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
por: Malek, Alan, et al.
Publicado: (2025)
por: Malek, Alan, et al.
Publicado: (2025)
Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
por: Fan, Chongyu, et al.
Publicado: (2024)
por: Fan, Chongyu, et al.
Publicado: (2024)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
por: Zheng, Xinzhe, et al.
Publicado: (2025)
por: Zheng, Xinzhe, et al.
Publicado: (2025)
Feature Subset Weighting for Distance-based Supervised Learning through Choquet Integration
por: Theerens, Adnan, et al.
Publicado: (2025)
por: Theerens, Adnan, et al.
Publicado: (2025)
Improving Rule-based Reasoning in LLMs using Neurosymbolic Representations
por: Dhanraj, Varun, et al.
Publicado: (2025)
por: Dhanraj, Varun, et al.
Publicado: (2025)
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
por: Gupta, Sharut, et al.
Publicado: (2026)
por: Gupta, Sharut, et al.
Publicado: (2026)
Efficiently Deploying LLMs with Controlled Risk
por: Zellinger, Michael J., et al.
Publicado: (2024)
por: Zellinger, Michael J., et al.
Publicado: (2024)
Physics-model-guided Worst-case Sampling for Safe Reinforcement Learning
por: Cao, Hongpeng, et al.
Publicado: (2024)
por: Cao, Hongpeng, et al.
Publicado: (2024)
Weight-based Decomposition: A Case for Bilinear MLPs
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Ejemplares similares
-
Differentially Private Worst-group Risk Minimization
por: Zhou, Xinyu, et al.
Publicado: (2024) -
Tamper-Resistant Safeguards for Open-Weight LLMs
por: Tamirisa, Rishub, et al.
Publicado: (2024) -
How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness
por: Rossolini, Giulio
Publicado: (2026) -
Open Problems in Frontier AI Risk Management
por: Ziosi, Marta, et al.
Publicado: (2026) -
Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
por: Aryal, Manish, et al.
Publicado: (2026)