Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.03056 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917244340011008 |
|---|---|
| author | Ren, Lu She, Junda Luo, Xinchen Wang, Tao Ye, Xin Zhang, Xu Wang, Muxuan Yang, Xiao Wang, Chenguang Xie, Fei Zhou, Yiwei Wu, Danjun Zhang, Guodong Hu, Yifei Zheng, Guoying Yang, Shujie Wang, Xingmei Wang, Shiyao Zhou, Yukun Yang, Fan Li, Size Cai, Kuo Luo, Qiang Tang, Ruiming Li, Han Gai, Kun |
| author_facet | Ren, Lu She, Junda Luo, Xinchen Wang, Tao Ye, Xin Zhang, Xu Wang, Muxuan Yang, Xiao Wang, Chenguang Xie, Fei Zhou, Yiwei Wu, Danjun Zhang, Guodong Hu, Yifei Zheng, Guoying Yang, Shujie Wang, Xingmei Wang, Shiyao Zhou, Yukun Yang, Fan Li, Size Cai, Kuo Luo, Qiang Tang, Ruiming Li, Han Gai, Kun |
| contents | Recent advances in large language models have highlighted their potential for personalized recommendation, where accurately capturing user preferences remains a key challenge. Leveraging their strong reasoning and generalization capabilities, LLMs offer new opportunities for modeling long-term user behavior. To systematically evaluate this, we introduce ALPBench, a Benchmark for Attribution-level Long-term Personal Behavior Understanding. Unlike item-focused benchmarks, ALPBench predicts user-interested attribute combinations, enabling ground-truth evaluation even for newly introduced items. It models preferences from long-term historical behaviors rather than users' explicitly expressed requests, better reflecting enduring interests. User histories are represented as natural language sequences, allowing interpretable, reasoning-based personalization. ALPBench enables fine-grained evaluation of personalization by focusing on the prediction of attribute combinations task that remains highly challenging for current LLMs due to the need to capture complex interactions among multiple attributes and reason over long-term user behavior sequences. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_03056 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ALPBench: A Benchmark for Attribution-level Long-term Personal Behavior Understanding Ren, Lu She, Junda Luo, Xinchen Wang, Tao Ye, Xin Zhang, Xu Wang, Muxuan Yang, Xiao Wang, Chenguang Xie, Fei Zhou, Yiwei Wu, Danjun Zhang, Guodong Hu, Yifei Zheng, Guoying Yang, Shujie Wang, Xingmei Wang, Shiyao Zhou, Yukun Yang, Fan Li, Size Cai, Kuo Luo, Qiang Tang, Ruiming Li, Han Gai, Kun Information Retrieval Recent advances in large language models have highlighted their potential for personalized recommendation, where accurately capturing user preferences remains a key challenge. Leveraging their strong reasoning and generalization capabilities, LLMs offer new opportunities for modeling long-term user behavior. To systematically evaluate this, we introduce ALPBench, a Benchmark for Attribution-level Long-term Personal Behavior Understanding. Unlike item-focused benchmarks, ALPBench predicts user-interested attribute combinations, enabling ground-truth evaluation even for newly introduced items. It models preferences from long-term historical behaviors rather than users' explicitly expressed requests, better reflecting enduring interests. User histories are represented as natural language sequences, allowing interpretable, reasoning-based personalization. ALPBench enables fine-grained evaluation of personalization by focusing on the prediction of attribute combinations task that remains highly challenging for current LLMs due to the need to capture complex interactions among multiple attributes and reason over long-term user behavior sequences. |
| title | ALPBench: A Benchmark for Attribution-level Long-term Personal Behavior Understanding |
| topic | Information Retrieval |
| url | https://arxiv.org/abs/2602.03056 |