AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yao, Jing, Duan, Shitong, Yi, Xiaoyuan, Xu, Dongkuan, Zhang, Peng, Lu, Tun, Gu, Ning, Dou, Zhicheng, Xie, Xing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915838753243136
author Yao, Jing
Duan, Shitong
Yi, Xiaoyuan
Xu, Dongkuan
Zhang, Peng
Lu, Tun
Gu, Ning
Dou, Zhicheng
Xie, Xing
author_facet Yao, Jing
Duan, Shitong
Yi, Xiaoyuan
Xu, Dongkuan
Zhang, Peng
Lu, Tun
Gu, Ning
Dou, Zhicheng
Xie, Xing
contents Assessing Large Language Models'(LLMs) underlying value differences enables comprehensive comparison of their misalignment, cultural adaptability, and biases. Nevertheless, current value measurement methods face the informativeness challenge: with often outdated, contaminated, or generic test questions, they can only capture the orientations on comment safety values, e.g., HHH, shared among different LLMs, leading to indistinguishable and uninformative results. To address this problem, we introduce AdAEM, a novel, self-extensible evaluation algorithm for revealing LLMs' inclinations. Distinct from static benchmarks, AdAEM automatically and adaptively generates and extends its test questions. This is achieved by probing the internal value boundaries of a diverse set of LLMs developed across cultures and time periods in an in-context optimization manner. Such a process theoretically maximizes an information-theoretic objective to extract diverse controversial topics that can provide more distinguishable and informative insights about models' value differences. In this way, AdAEM is able to co-evolve with the development of LLMs, consistently tracking their value dynamics. We use AdAEM to generate novel questions and conduct an extensive analysis, demonstrating our method's validity and effectiveness, laying the groundwork for better interdisciplinary research on LLMs' values and alignment. Codes and the generated evaluation questions are released at https://github.com/ValueCompass/AdAEM.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13531
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
Yao, Jing
Duan, Shitong
Yi, Xiaoyuan
Xu, Dongkuan
Zhang, Peng
Lu, Tun
Gu, Ning
Dou, Zhicheng
Xie, Xing
Computers and Society
Artificial Intelligence
Computation and Language
Assessing Large Language Models'(LLMs) underlying value differences enables comprehensive comparison of their misalignment, cultural adaptability, and biases. Nevertheless, current value measurement methods face the informativeness challenge: with often outdated, contaminated, or generic test questions, they can only capture the orientations on comment safety values, e.g., HHH, shared among different LLMs, leading to indistinguishable and uninformative results. To address this problem, we introduce AdAEM, a novel, self-extensible evaluation algorithm for revealing LLMs' inclinations. Distinct from static benchmarks, AdAEM automatically and adaptively generates and extends its test questions. This is achieved by probing the internal value boundaries of a diverse set of LLMs developed across cultures and time periods in an in-context optimization manner. Such a process theoretically maximizes an information-theoretic objective to extract diverse controversial topics that can provide more distinguishable and informative insights about models' value differences. In this way, AdAEM is able to co-evolve with the development of LLMs, consistently tracking their value dynamics. We use AdAEM to generate novel questions and conduct an extensive analysis, demonstrating our method's validity and effectiveness, laying the groundwork for better interdisciplinary research on LLMs' values and alignment. Codes and the generated evaluation questions are released at https://github.com/ValueCompass/AdAEM.
title AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
topic Computers and Society
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.13531