Automated alignment is harder than you think
Fuente:
arXiv
Salvato in:
| Autori principali: | Bowkis, Aleksandr, Buhl, Marie Davidsen, Pfau, Jacob, Irving, Geoffrey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An alignment safety case sketch based on debate
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
Emerging Practices in Frontier AI Safety Frameworks
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
di: Oliva, Maria Paz, et al.
Pubblicazione: (2025)
di: Oliva, Maria Paz, et al.
Pubblicazione: (2025)
The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
di: Shostack, Adam
Pubblicazione: (2024)
di: Shostack, Adam
Pubblicazione: (2024)
Before you <think>, monitor: Implementing Flavell's metacognitive framework in LLMs
di: Oh, Nick
Pubblicazione: (2025)
di: Oh, Nick
Pubblicazione: (2025)
Your device may know you better than you know yourself -- continuous authentication on novel dataset using machine learning
di: Nascimento, Pedro Gomes do, et al.
Pubblicazione: (2024)
di: Nascimento, Pedro Gomes do, et al.
Pubblicazione: (2024)
MEANT: Multimodal Encoder for Antecedent Information
di: Irving, Benjamin Iyoya, et al.
Pubblicazione: (2024)
di: Irving, Benjamin Iyoya, et al.
Pubblicazione: (2024)
Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
di: Pfau, Jacob, et al.
Pubblicazione: (2024)
di: Pfau, Jacob, et al.
Pubblicazione: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
Safety case template for frontier AI: A cyber inability argument
di: Goemans, Arthur, et al.
Pubblicazione: (2024)
di: Goemans, Arthur, et al.
Pubblicazione: (2024)
AlphaBeta is not as good as you think: a simple class of synthetic games for a better analysis of deterministic game-solving algorithms
di: Boige, Raphaël, et al.
Pubblicazione: (2025)
di: Boige, Raphaël, et al.
Pubblicazione: (2025)
From LLM-Driven Trading Card Generation to Procedural Relatedness: A Pokémon Case Study
di: Pfau, Johannes, et al.
Pubblicazione: (2026)
di: Pfau, Johannes, et al.
Pubblicazione: (2026)
Avoiding Obfuscation with Prover-Estimator Debate
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2025)
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2025)
Engineering Trustworthy AI: A Developer Guide for Empirical Risk Minimization
di: Pfau, Diana, et al.
Pubblicazione: (2024)
di: Pfau, Diana, et al.
Pubblicazione: (2024)
OntView: What you See is What you Meant
di: Bobed, Carlos, et al.
Pubblicazione: (2025)
di: Bobed, Carlos, et al.
Pubblicazione: (2025)
Human-aligned Chess with a Bit of Search
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
A sketch of an AI control safety case
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
di: Korbak, Tomek, et al.
Pubblicazione: (2025)
How critically can an AI think? A framework for evaluating the quality of thinking of generative artificial intelligence
di: Zaphir, Luke, et al.
Pubblicazione: (2024)
di: Zaphir, Luke, et al.
Pubblicazione: (2024)
What you get is what you see: Decomposing Epistemic Planning using Functional STRIPS
di: Hu, Guang, et al.
Pubblicazione: (2019)
di: Hu, Guang, et al.
Pubblicazione: (2019)
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
di: Meincke, Lennart, et al.
Pubblicazione: (2025)
di: Meincke, Lennart, et al.
Pubblicazione: (2025)
Explainers' Mental Representations of Explainees' Needs in Everyday Explanations
di: Schaffer, Michael Erol, et al.
Pubblicazione: (2024)
di: Schaffer, Michael Erol, et al.
Pubblicazione: (2024)
PPSZ is better than you think
di: Scheder, Dominik
Pubblicazione: (2022)
di: Scheder, Dominik
Pubblicazione: (2022)
Debate is efficient with your time
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
di: Pathmanathan, Pankayaraj, et al.
Pubblicazione: (2024)
di: Pathmanathan, Pankayaraj, et al.
Pubblicazione: (2024)
'Too much alignment; not enough culture': Re-balancing cultural alignment practices in LLMs
di: Orlowski, Eric J. W., et al.
Pubblicazione: (2025)
di: Orlowski, Eric J. W., et al.
Pubblicazione: (2025)
REL: Working out is all you need
di: Simonds, Toby, et al.
Pubblicazione: (2024)
di: Simonds, Toby, et al.
Pubblicazione: (2024)
Extracting alignment data in open models
di: Barbero, Federico, et al.
Pubblicazione: (2025)
di: Barbero, Federico, et al.
Pubblicazione: (2025)
Value alignment: a formal approach
di: Sierra, Carles, et al.
Pubblicazione: (2021)
di: Sierra, Carles, et al.
Pubblicazione: (2021)
Segmentation Re-thinking Uncertainty Estimation Metrics for Semantic Segmentation
di: Ma, Qitian, et al.
Pubblicazione: (2024)
di: Ma, Qitian, et al.
Pubblicazione: (2024)
Scaling up the think-aloud method
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
Do not think about pink elephant!
di: Hwang, Kyomin, et al.
Pubblicazione: (2024)
di: Hwang, Kyomin, et al.
Pubblicazione: (2024)
Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean
di: Meshkov, Aleksandr
Pubblicazione: (2026)
di: Meshkov, Aleksandr
Pubblicazione: (2026)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
di: Pletenev, Sergey, et al.
Pubblicazione: (2025)
di: Pletenev, Sergey, et al.
Pubblicazione: (2025)
Attention when you need
di: Boominathan, Lokesh, et al.
Pubblicazione: (2025)
di: Boominathan, Lokesh, et al.
Pubblicazione: (2025)
Propaganda is all you need
di: Kronlund-Drouault, Paul
Pubblicazione: (2024)
di: Kronlund-Drouault, Paul
Pubblicazione: (2024)
Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
di: Luyen, Le Ngoc, et al.
Pubblicazione: (2025)
di: Luyen, Le Ngoc, et al.
Pubblicazione: (2025)
Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation?
di: Li, Wensu, et al.
Pubblicazione: (2026)
di: Li, Wensu, et al.
Pubblicazione: (2026)
Compression is all you need: Modeling Mathematics
di: Aksenov, Vitaly, et al.
Pubblicazione: (2026)
di: Aksenov, Vitaly, et al.
Pubblicazione: (2026)
Practical challenges of control monitoring in frontier AI deployments
di: Lindner, David, et al.
Pubblicazione: (2025)
di: Lindner, David, et al.
Pubblicazione: (2025)
Documenti analoghi
-
An alignment safety case sketch based on debate
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025) -
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025) -
Emerging Practices in Frontier AI Safety Frameworks
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025) -
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
di: Oliva, Maria Paz, et al.
Pubblicazione: (2025) -
The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
di: Shostack, Adam
Pubblicazione: (2024)