Evaluating Agents using Social Choice Theory
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918073845415936 |
|---|---|
| author | Lanctot, Marc Larson, Kate Bachrach, Yoram Marris, Luke Li, Zun Bhoopchand, Avishkar Anthony, Thomas Tanner, Brian Koop, Anna |
| author_facet | Lanctot, Marc Larson, Kate Bachrach, Yoram Marris, Luke Li, Zun Bhoopchand, Avishkar Anthony, Thomas Tanner, Brian Koop, Anna |
| contents | We argue that many general evaluation problems can be viewed through the lens of voting theory. Each task is interpreted as a separate voter, which requires only ordinal rankings or pairwise comparisons of agents to produce an overall evaluation. By viewing the aggregator as a social welfare function, we are able to leverage centuries of research in social choice theory to derive principled evaluation frameworks with axiomatic foundations. These evaluations are interpretable and flexible, while avoiding many of the problems currently facing cross-task evaluation. We apply this Voting-as-Evaluation (VasE) framework across multiple settings, including reinforcement learning, large language models, and humans. In practice, we observe that VasE can be more robust than popular evaluation frameworks (Elo and Nash averaging), discovers properties in the evaluation data not evident from scores alone, and can predict outcomes better than Elo in a complex seven-player game. We identify one particular approach, maximal lotteries, that satisfies important consistency properties relevant to evaluation, is computationally efficient (polynomial in the size of the evaluation data), and identifies game-theoretic cycles. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_03121 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Evaluating Agents using Social Choice Theory Lanctot, Marc Larson, Kate Bachrach, Yoram Marris, Luke Li, Zun Bhoopchand, Avishkar Anthony, Thomas Tanner, Brian Koop, Anna Artificial Intelligence Computer Science and Game Theory Multiagent Systems We argue that many general evaluation problems can be viewed through the lens of voting theory. Each task is interpreted as a separate voter, which requires only ordinal rankings or pairwise comparisons of agents to produce an overall evaluation. By viewing the aggregator as a social welfare function, we are able to leverage centuries of research in social choice theory to derive principled evaluation frameworks with axiomatic foundations. These evaluations are interpretable and flexible, while avoiding many of the problems currently facing cross-task evaluation. We apply this Voting-as-Evaluation (VasE) framework across multiple settings, including reinforcement learning, large language models, and humans. In practice, we observe that VasE can be more robust than popular evaluation frameworks (Elo and Nash averaging), discovers properties in the evaluation data not evident from scores alone, and can predict outcomes better than Elo in a complex seven-player game. We identify one particular approach, maximal lotteries, that satisfies important consistency properties relevant to evaluation, is computationally efficient (polynomial in the size of the evaluation data), and identifies game-theoretic cycles. |
| title | Evaluating Agents using Social Choice Theory |
| topic | Artificial Intelligence Computer Science and Game Theory Multiagent Systems |
| url | https://arxiv.org/abs/2312.03121 |