BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pihlakas, Roland, Kuriakose, Sruthi Susan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks
von: Pihlakas, Roland
Veröffentlicht: (2024)
von: Pihlakas, Roland
Veröffentlicht: (2024)
levitation-opensource/Manipulative-Expression-Recognition: v1.5.12
von: Roland Pihlakas
Veröffentlicht: (2026)
von: Roland Pihlakas
Veröffentlicht: (2026)
levitation-opensource/MboxComparer: v1.2.1
von: Roland Pihlakas
Veröffentlicht: (2026)
von: Roland Pihlakas
Veröffentlicht: (2026)
Probabilistic Calibration Is a Trainable Capability in Language Models
von: Baldelli, Davide, et al.
Veröffentlicht: (2026)
von: Baldelli, Davide, et al.
Veröffentlicht: (2026)
Test-time RL alignment exposes task familiarity artifacts in LLM benchmarks
von: Wang, Kun, et al.
Veröffentlicht: (2026)
von: Wang, Kun, et al.
Veröffentlicht: (2026)
Annotation alignment: Comparing LLM and human annotations of conversational safety
von: Movva, Rajiv, et al.
Veröffentlicht: (2024)
von: Movva, Rajiv, et al.
Veröffentlicht: (2024)
Does optimisation of heart failure equate to fitness for surgery?
von: Sebastian Vaughan‐Burleigh, et al.
Veröffentlicht: (2025)
von: Sebastian Vaughan‐Burleigh, et al.
Veröffentlicht: (2025)
High availability for cloud-based active–active services
von: Kuriakose John, Linton
Veröffentlicht: (2025)
von: Kuriakose John, Linton
Veröffentlicht: (2025)
Prioritise safety, optimise success! Return to rugby postpartum
von: GM Donnelly, et al.
Veröffentlicht: (2024)
von: GM Donnelly, et al.
Veröffentlicht: (2024)
MoEITS: A Green AI approach for simplifying MoE-LLMs
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
The observable impact of runaway OB stars on protoplanetary discs
von: Coleman, Gavin A. L., et al.
Veröffentlicht: (2025)
von: Coleman, Gavin A. L., et al.
Veröffentlicht: (2025)
Discovery of a runaway star likely ejected by a Type Iax Supernova
von: Bhat, A., et al.
Veröffentlicht: (2026)
von: Bhat, A., et al.
Veröffentlicht: (2026)
The formation of inhomogeneities in gravitational string like mode
von: Kozlov, G. A.
Veröffentlicht: (2026)
von: Kozlov, G. A.
Veröffentlicht: (2026)
Chance constraints transcription and failure risk estimation for stochastic trajectory optimisation
von: Caleb, Thomas, et al.
Veröffentlicht: (2025)
von: Caleb, Thomas, et al.
Veröffentlicht: (2025)
The puzzling failure of economics
Veröffentlicht: (1997)
Veröffentlicht: (1997)
Usefulness of failure mode and effects analysis for improving mobilization safety in critically ill patients
von: Agustín Vázquez-Valencia
Veröffentlicht: (2018)
von: Agustín Vázquez-Valencia
Veröffentlicht: (2018)
A simplified characterization of stable-like heat kernel estimates
von: Murugan, Mathav
Veröffentlicht: (2026)
von: Murugan, Mathav
Veröffentlicht: (2026)
Perspectives on benchmarking foundation models for network biology
von: Christina V. Theodoris
Veröffentlicht: (2024)
von: Christina V. Theodoris
Veröffentlicht: (2024)
An observational study of rotation and binarity of Galactic O-type runaway stars
von: Carretero-Castrillo, M., et al.
Veröffentlicht: (2025)
von: Carretero-Castrillo, M., et al.
Veröffentlicht: (2025)
Systematic improvement of the quantum approximate optimisation ansatz for combinatorial optimisation using quantum subspace expansion
von: Beaujeault-Taudière, Yann
Veröffentlicht: (2025)
von: Beaujeault-Taudière, Yann
Veröffentlicht: (2025)
OAEI 2004 (Ontology alignment contest) benchmark tests (oacontest09)
von: Euzenat, Jérôme
Veröffentlicht: (2004)
von: Euzenat, Jérôme
Veröffentlicht: (2004)
Identifying close-in Jupiters that arrived via disk migration: Evidence of primordial alignment, preference of nearby companions and hint of runaway migration
von: Kawai, Yugo, et al.
Veröffentlicht: (2025)
von: Kawai, Yugo, et al.
Veröffentlicht: (2025)
An alignment safety case sketch based on debate
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2025)
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2025)
The economic alignment problem of artificial intelligence
von: O'Neill, Daniel W., et al.
Veröffentlicht: (2026)
von: O'Neill, Daniel W., et al.
Veröffentlicht: (2026)
What shapes the formation of interstate benchmarking networks?
von: Shuai Cao, et al.
Veröffentlicht: (2024)
von: Shuai Cao, et al.
Veröffentlicht: (2024)
Stellar control on atmospheric carbon chemistry, CO runaway, and organic synthesis on lifeless Earth-like planets
von: Endo, Yoshiaki, et al.
Veröffentlicht: (2026)
von: Endo, Yoshiaki, et al.
Veröffentlicht: (2026)
Unionizing the Ivory Tower: Cornell workers’ fifteen‐year fight for justice and a living wage By AlDavidoff (2023). Ithaca and London: ILR Press. 238 pages, ISBN: 9781501771552
von: Deepa Kylasam Iyer, et al.
Veröffentlicht: (2024)
von: Deepa Kylasam Iyer, et al.
Veröffentlicht: (2024)
Minor Keys: Gender, Inequality and Work in Electronic Music
von: Deepa Kylasam Iyer, et al.
Veröffentlicht: (2026)
von: Deepa Kylasam Iyer, et al.
Veröffentlicht: (2026)
Orbit-averaging and deposition accuracy for runaway electron beams in hybrid kinetic-MHD simulations of the runaway plateau
von: López, O. E., et al.
Veröffentlicht: (2025)
von: López, O. E., et al.
Veröffentlicht: (2025)
What should an AI assessor optimise for?
von: Romero-Alvarado, Daniel, et al.
Veröffentlicht: (2025)
von: Romero-Alvarado, Daniel, et al.
Veröffentlicht: (2025)
Static native tibial alignment in total knee arthroplasty optimises whole‐body gait kinematics
von: Zhijun Li, et al.
Veröffentlicht: (2026)
von: Zhijun Li, et al.
Veröffentlicht: (2026)
Stability of Equatorial modes in a simplified coupled ocean-atmosphere model
von: Chunzai Wang and Robert H. Weisberg
Veröffentlicht: (1996)
von: Chunzai Wang and Robert H. Weisberg
Veröffentlicht: (1996)
Actionable AI: Enabling Non Experts to Understand and Configure AI Systems
von: Boulard, Cécile, et al.
Veröffentlicht: (2025)
von: Boulard, Cécile, et al.
Veröffentlicht: (2025)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
von: Kankowski, Florian, et al.
Veröffentlicht: (2025)
von: Kankowski, Florian, et al.
Veröffentlicht: (2025)
Inverse Design of Inorganic Compounds with Generative AI
von: Kneiding, Hannes, et al.
Veröffentlicht: (2026)
von: Kneiding, Hannes, et al.
Veröffentlicht: (2026)
Colour difference and prediction optimisation of silk‐like plain knitted fabrics
von: Shuqi Huang, et al.
Veröffentlicht: (2025)
von: Shuqi Huang, et al.
Veröffentlicht: (2025)
BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists
von: Provatas, Kimon Antonios, et al.
Veröffentlicht: (2026)
von: Provatas, Kimon Antonios, et al.
Veröffentlicht: (2026)
Marginality from Leading Soft Gluons
von: Narayanan, Sruthi A.
Veröffentlicht: (2024)
von: Narayanan, Sruthi A.
Veröffentlicht: (2024)
FROST-CLUSTERS -- III. Metallicity-dependent intermediate mass black hole formation by runaway collisions in dense star clusters
von: Rantala, Antti, et al.
Veröffentlicht: (2026)
von: Rantala, Antti, et al.
Veröffentlicht: (2026)
Evaluation format, not model capability, drives triage failure in the assessment of consumer health AI
von: Navarro, David Fraile, et al.
Veröffentlicht: (2026)
von: Navarro, David Fraile, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks
von: Pihlakas, Roland
Veröffentlicht: (2024) -
levitation-opensource/Manipulative-Expression-Recognition: v1.5.12
von: Roland Pihlakas
Veröffentlicht: (2026) -
levitation-opensource/MboxComparer: v1.2.1
von: Roland Pihlakas
Veröffentlicht: (2026) -
Probabilistic Calibration Is a Trainable Capability in Language Models
von: Baldelli, Davide, et al.
Veröffentlicht: (2026) -
Test-time RL alignment exposes task familiarity artifacts in LLM benchmarks
von: Wang, Kun, et al.
Veröffentlicht: (2026)