CharacterBench: Benchmarking Character Customization of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915066130989056 |
|---|---|
| author | Zhou, Jinfeng Huang, Yongkang Wen, Bosi Bi, Guanqun Chen, Yuxuan Ke, Pei Chen, Zhuang Xiao, Xiyao Peng, Libiao Tang, Kuntian Zhang, Rongsheng Zhang, Le Lv, Tangjie Hu, Zhipeng Wang, Hongning Huang, Minlie |
| author_facet | Zhou, Jinfeng Huang, Yongkang Wen, Bosi Bi, Guanqun Chen, Yuxuan Ke, Pei Chen, Zhuang Xiao, Xiyao Peng, Libiao Tang, Kuntian Zhang, Rongsheng Zhang, Le Lv, Tangjie Hu, Zhipeng Wang, Hongning Huang, Minlie |
| contents | Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs' character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters' responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark's potential to optimize LLMs' character customization. Our repository is at https://github.com/thu-coai/CharacterBench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_11912 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | CharacterBench: Benchmarking Character Customization of Large Language Models Zhou, Jinfeng Huang, Yongkang Wen, Bosi Bi, Guanqun Chen, Yuxuan Ke, Pei Chen, Zhuang Xiao, Xiyao Peng, Libiao Tang, Kuntian Zhang, Rongsheng Zhang, Le Lv, Tangjie Hu, Zhipeng Wang, Hongning Huang, Minlie Computation and Language Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs' character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters' responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark's potential to optimize LLMs' character customization. Our repository is at https://github.com/thu-coai/CharacterBench. |
| title | CharacterBench: Benchmarking Character Customization of Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.11912 |