A Revealed Preference Framework for AI Alignment
Fuente:
arXiv
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914431368167424 |
|---|---|
| author | Suleymanov, Elchin |
| author_facet | Suleymanov, Elchin |
| contents | Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal's preferences or pursue its own? To study this question using revealed preference techniques, I introduce the Luce Alignment Model, where the AI's choices are a mixture of two Luce rules, one reflecting the human's preferences and the other the AI's. I show that the AI's alignment (similarity of human and AI preferences) can be generically identified in two settings: the laboratory setting, where both human and AI choices are observed, and the field setting, where only AI choices are observed. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_27868 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | A Revealed Preference Framework for AI Alignment Suleymanov, Elchin Theoretical Economics Artificial Intelligence Computer Science and Game Theory Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal's preferences or pursue its own? To study this question using revealed preference techniques, I introduce the Luce Alignment Model, where the AI's choices are a mixture of two Luce rules, one reflecting the human's preferences and the other the AI's. I show that the AI's alignment (similarity of human and AI preferences) can be generically identified in two settings: the laboratory setting, where both human and AI choices are observed, and the field setting, where only AI choices are observed. |
| title | A Revealed Preference Framework for AI Alignment |
| topic | Theoretical Economics Artificial Intelligence Computer Science and Game Theory |
| url | https://arxiv.org/abs/2603.27868 |