CL-bench Life: Can Language Models Learn from Real-Life Context?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dou, Shihan, Shen, Yujiong, Huang, Chenhao, Ye, Junjie, Chen, Jiayi, Wang, Junzhe, He, Qianyu, Liu, Shichun, Lv, Changze, Lin, Jiahang, Zhang, Jiazheng, Zhang, Ming, Liu, Shaofan, Ji, Tao, Yin, Zhangyue, Zhang, Cheng, Xie, Huaibing, Hu, Jianglu, Deng, Jingcheng, Li, Lincheng, Hu, Minda, Wang, Shaolei, Zhao, Syrus, Wang, Weichao, Lei, Yan, Liu, Yang, Xiao, Yanling, Liu, Yiting, Xu, Zenan, Guo, Zhen, Zhao, Ziliang, Zhou, Pluto, Gui, Tao, Zhang, Qi, Huang, Xuanjing, Jiang, Yu-Gang, Wang, Di, Yao, Shunyu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918474841849856
author Dou, Shihan
Shen, Yujiong
Huang, Chenhao
Ye, Junjie
Chen, Jiayi
Wang, Junzhe
He, Qianyu
Liu, Shichun
Lv, Changze
Lin, Jiahang
Zhang, Jiazheng
Zhang, Ming
Liu, Shaofan
Ji, Tao
Yin, Zhangyue
Zhang, Cheng
Xie, Huaibing
Hu, Jianglu
Deng, Jingcheng
Li, Lincheng
Hu, Minda
Wang, Shaolei
Zhao, Syrus
Wang, Weichao
Lei, Yan
Liu, Yang
Xiao, Yanling
Liu, Yiting
Xu, Zenan
Guo, Zhen
Zhao, Ziliang
Zhou, Pluto
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Jiang, Yu-Gang
Wang, Di
Yao, Shunyu
author_facet Dou, Shihan
Shen, Yujiong
Huang, Chenhao
Ye, Junjie
Chen, Jiayi
Wang, Junzhe
He, Qianyu
Liu, Shichun
Lv, Changze
Lin, Jiahang
Zhang, Jiazheng
Zhang, Ming
Liu, Shaofan
Ji, Tao
Yin, Zhangyue
Zhang, Cheng
Xie, Huaibing
Hu, Jianglu
Deng, Jingcheng
Li, Lincheng
Hu, Minda
Wang, Shaolei
Zhao, Syrus
Wang, Weichao
Lei, Yan
Liu, Yang
Xiao, Yanling
Liu, Yiting
Xu, Zenan
Guo, Zhen
Zhao, Ziliang
Zhou, Pluto
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Jiang, Yu-Gang
Wang, Di
Yao, Shunyu
contents Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for models. As these systems move beyond professional settings into everyday life, the nature of the contexts they must handle also shifts. Real-life contexts are often messy, fragmented, and deeply tied to personal and social experience, such as multi-party conversations, personal archives, and behavioral traces. Yet it remains unclear whether current frontier language models can reliably learn from such contexts and solve tasks grounded in them. To this end, we introduce CL-bench Life, a fully human-curated benchmark comprising 405 context-task pairs and 5,348 verification rubrics, covering common real-life scenarios. Solving tasks in CL-bench Life requires models to reason over complex, messy real-life contexts, calling for strong real-life context learning abilities that go far beyond those evaluated in existing benchmarks. We evaluate ten frontier LMs and find that real-life context learning remains highly challenging: even the best-performing model achieves only 19.3% task solving rate, while the average performance across models is only 13.8%. Models still struggle to reason over contexts such as messy group chat histories and fragmented behavioral records from everyday life. CL-bench Life provides a crucial testbed for advancing real-life context learning, and progress on it can enable more intelligent and reliable AI assistants in everyday life.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27043
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CL-bench Life: Can Language Models Learn from Real-Life Context?
Dou, Shihan
Shen, Yujiong
Huang, Chenhao
Ye, Junjie
Chen, Jiayi
Wang, Junzhe
He, Qianyu
Liu, Shichun
Lv, Changze
Lin, Jiahang
Zhang, Jiazheng
Zhang, Ming
Liu, Shaofan
Ji, Tao
Yin, Zhangyue
Zhang, Cheng
Xie, Huaibing
Hu, Jianglu
Deng, Jingcheng
Li, Lincheng
Hu, Minda
Wang, Shaolei
Zhao, Syrus
Wang, Weichao
Lei, Yan
Liu, Yang
Xiao, Yanling
Liu, Yiting
Xu, Zenan
Guo, Zhen
Zhao, Ziliang
Zhou, Pluto
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Jiang, Yu-Gang
Wang, Di
Yao, Shunyu
Computation and Language
Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for models. As these systems move beyond professional settings into everyday life, the nature of the contexts they must handle also shifts. Real-life contexts are often messy, fragmented, and deeply tied to personal and social experience, such as multi-party conversations, personal archives, and behavioral traces. Yet it remains unclear whether current frontier language models can reliably learn from such contexts and solve tasks grounded in them. To this end, we introduce CL-bench Life, a fully human-curated benchmark comprising 405 context-task pairs and 5,348 verification rubrics, covering common real-life scenarios. Solving tasks in CL-bench Life requires models to reason over complex, messy real-life contexts, calling for strong real-life context learning abilities that go far beyond those evaluated in existing benchmarks. We evaluate ten frontier LMs and find that real-life context learning remains highly challenging: even the best-performing model achieves only 19.3% task solving rate, while the average performance across models is only 13.8%. Models still struggle to reason over contexts such as messy group chat histories and fragmented behavioral records from everyday life. CL-bench Life provides a crucial testbed for advancing real-life context learning, and progress on it can enable more intelligent and reliable AI assistants in everyday life.
title CL-bench Life: Can Language Models Learn from Real-Life Context?
topic Computation and Language
url https://arxiv.org/abs/2604.27043