An alignment safety case sketch based on debate
Fuente:
arXiv
Salvato in:
| Autori principali: | Buhl, Marie Davidsen, Pfau, Jacob, Hilton, Benjamin, Irving, Geoffrey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automated alignment is harder than you think
di: Bowkis, Aleksandr, et al.
Pubblicazione: (2026)
di: Bowkis, Aleksandr, et al.
Pubblicazione: (2026)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
A sketch of an AI control safety case
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
Emerging Practices in Frontier AI Safety Frameworks
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
Safety case template for frontier AI: A cyber inability argument
di: Goemans, Arthur, et al.
Pubblicazione: (2024)
di: Goemans, Arthur, et al.
Pubblicazione: (2024)
When can we trust untrusted monitoring? A safety case sketch across collusion strategies
di: Gardner-Challis, Nelson, et al.
Pubblicazione: (2026)
di: Gardner-Challis, Nelson, et al.
Pubblicazione: (2026)
MEANT: Multimodal Encoder for Antecedent Information
di: Irving, Benjamin Iyoya, et al.
Pubblicazione: (2024)
di: Irving, Benjamin Iyoya, et al.
Pubblicazione: (2024)
Estimating the Probabilities of Rare Outputs in Language Models
di: Wu, Gabriel, et al.
Pubblicazione: (2024)
di: Wu, Gabriel, et al.
Pubblicazione: (2024)
Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
di: Pfau, Jacob, et al.
Pubblicazione: (2024)
di: Pfau, Jacob, et al.
Pubblicazione: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
Towards evaluations-based safety cases for AI scheming
di: Balesni, Mikita, et al.
Pubblicazione: (2024)
di: Balesni, Mikita, et al.
Pubblicazione: (2024)
Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration
di: He, Yuxiang, et al.
Pubblicazione: (2025)
di: He, Yuxiang, et al.
Pubblicazione: (2025)
From LLM-Driven Trading Card Generation to Procedural Relatedness: A Pokémon Case Study
di: Pfau, Johannes, et al.
Pubblicazione: (2026)
di: Pfau, Johannes, et al.
Pubblicazione: (2026)
Avoiding Obfuscation with Prover-Estimator Debate
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2025)
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2025)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
di: Hu, Mengxuan, et al.
Pubblicazione: (2026)
di: Hu, Mengxuan, et al.
Pubblicazione: (2026)
From edges to meaning: Semantic line sketches as a cognitive scaffold for ancient pictograph invention
di: Leem, Seowung, et al.
Pubblicazione: (2026)
di: Leem, Seowung, et al.
Pubblicazione: (2026)
From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks
di: Pihlakas, Roland
Pubblicazione: (2024)
di: Pihlakas, Roland
Pubblicazione: (2024)
A Subjective Logic-based method for runtime confidence updates in safety arguments
di: Herd, Benjamin, et al.
Pubblicazione: (2026)
di: Herd, Benjamin, et al.
Pubblicazione: (2026)
Engineering Trustworthy AI: A Developer Guide for Empirical Risk Minimization
di: Pfau, Diana, et al.
Pubblicazione: (2024)
di: Pfau, Diana, et al.
Pubblicazione: (2024)
Vector sketch animation generation with differentiable motion trajectories
di: Zhu, Xinding, et al.
Pubblicazione: (2025)
di: Zhu, Xinding, et al.
Pubblicazione: (2025)
Human-aligned Chess with a Bit of Search
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
Safety cases for frontier AI
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2024)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2024)
Backdoor defense, learnability and obfuscation
di: Christiano, Paul, et al.
Pubblicazione: (2024)
di: Christiano, Paul, et al.
Pubblicazione: (2024)
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
Towards a Law of Iterated Expectations for Heuristic Estimators
di: Christiano, Paul, et al.
Pubblicazione: (2024)
di: Christiano, Paul, et al.
Pubblicazione: (2024)
Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation
di: Li, Xinshu, et al.
Pubblicazione: (2025)
di: Li, Xinshu, et al.
Pubblicazione: (2025)
Explainers' Mental Representations of Explainees' Needs in Everyday Explanations
di: Schaffer, Michael Erol, et al.
Pubblicazione: (2024)
di: Schaffer, Michael Erol, et al.
Pubblicazione: (2024)
Debate is efficient with your time
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
'Too much alignment; not enough culture': Re-balancing cultural alignment practices in LLMs
di: Orlowski, Eric J. W., et al.
Pubblicazione: (2025)
di: Orlowski, Eric J. W., et al.
Pubblicazione: (2025)
Hybrid Classical-Quantum architecture for vectorised image classification of hand-written sketches
di: Cordero, Y., et al.
Pubblicazione: (2024)
di: Cordero, Y., et al.
Pubblicazione: (2024)
Extracting alignment data in open models
di: Barbero, Federico, et al.
Pubblicazione: (2025)
di: Barbero, Federico, et al.
Pubblicazione: (2025)
Value alignment: a formal approach
di: Sierra, Carles, et al.
Pubblicazione: (2021)
di: Sierra, Carles, et al.
Pubblicazione: (2021)
The impact of multi-agent debate protocols on debate quality: a controlled case study
di: Marandi, Ramtin Zargari
Pubblicazione: (2026)
di: Marandi, Ramtin Zargari
Pubblicazione: (2026)
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
di: Pihlakas, Roland, et al.
Pubblicazione: (2025)
di: Pihlakas, Roland, et al.
Pubblicazione: (2025)
MASER: Modality-Adaptive Specialist Routing for Embodied 3D Spatial Intelligence
di: Raj, Hilton, et al.
Pubblicazione: (2026)
di: Raj, Hilton, et al.
Pubblicazione: (2026)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
di: Shojaei, Mostafa Faghih, et al.
Pubblicazione: (2025)
di: Shojaei, Mostafa Faghih, et al.
Pubblicazione: (2025)
Landscape of AI safety concerns -- A methodology to support safety assurance for AI-based autonomous systems
di: Schnitzer, Ronald, et al.
Pubblicazione: (2024)
di: Schnitzer, Ronald, et al.
Pubblicazione: (2024)
Practical challenges of control monitoring in frontier AI deployments
di: Lindner, David, et al.
Pubblicazione: (2025)
di: Lindner, David, et al.
Pubblicazione: (2025)
Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs
di: Xuan-Vu, Nguyen, et al.
Pubblicazione: (2025)
di: Xuan-Vu, Nguyen, et al.
Pubblicazione: (2025)
Super Co-alignment of Human and AI for Sustainable Symbiotic Society
di: Zeng, Yi, et al.
Pubblicazione: (2025)
di: Zeng, Yi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Automated alignment is harder than you think
di: Bowkis, Aleksandr, et al.
Pubblicazione: (2026) -
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025) -
A sketch of an AI control safety case
di: Korbak, Tomek, et al.
Pubblicazione: (2025) -
Emerging Practices in Frontier AI Safety Frameworks
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025) -
Safety case template for frontier AI: A cyber inability argument
di: Goemans, Arthur, et al.
Pubblicazione: (2024)