CL-bench: A Benchmark for Context Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dou, Shihan, Zhang, Ming, Yin, Zhangyue, Huang, Chenhao, Shen, Yujiong, Wang, Junzhe, Chen, Jiayi, Ni, Yuchen, Ye, Junjie, Zhang, Cheng, Xie, Huaibing, Hu, Jianglu, Wang, Shaolei, Wang, Weichao, Xiao, Yanling, Liu, Yiting, Xu, Zenan, Guo, Zhen, Zhou, Pluto, Gui, Tao, Wu, Zuxuan, Qiu, Xipeng, Zhang, Qi, Huang, Xuanjing, Jiang, Yu-Gang, Wang, Di, Yao, Shunyu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915771411595264
author Dou, Shihan
Zhang, Ming
Yin, Zhangyue
Huang, Chenhao
Shen, Yujiong
Wang, Junzhe
Chen, Jiayi
Ni, Yuchen
Ye, Junjie
Zhang, Cheng
Xie, Huaibing
Hu, Jianglu
Wang, Shaolei
Wang, Weichao
Xiao, Yanling
Liu, Yiting
Xu, Zenan
Guo, Zhen
Zhou, Pluto
Gui, Tao
Wu, Zuxuan
Qiu, Xipeng
Zhang, Qi
Huang, Xuanjing
Jiang, Yu-Gang
Wang, Di
Yao, Shunyu
author_facet Dou, Shihan
Zhang, Ming
Yin, Zhangyue
Huang, Chenhao
Shen, Yujiong
Wang, Junzhe
Chen, Jiayi
Ni, Yuchen
Ye, Junjie
Zhang, Cheng
Xie, Huaibing
Hu, Jianglu
Wang, Shaolei
Wang, Weichao
Xiao, Yanling
Liu, Yiting
Xu, Zenan
Guo, Zhen
Zhou, Pluto
Gui, Tao
Wu, Zuxuan
Qiu, Xipeng
Zhang, Qi
Huang, Xuanjing
Jiang, Yu-Gang
Wang, Di
Yao, Shunyu
contents Current language models (LMs) excel at reasoning over prompts using pre-trained knowledge. However, real-world tasks are far more complex and context-dependent: models must learn from task-specific context and leverage new knowledge beyond what is learned during pre-training to reason and resolve tasks. We term this capability context learning, a crucial ability that humans naturally possess but has been largely overlooked. To this end, we introduce CL-bench, a real-world benchmark consisting of 500 complex contexts, 1,899 tasks, and 31,607 verification rubrics, all crafted by experienced domain experts. Each task is designed such that the new content required to resolve it is contained within the corresponding context. Resolving tasks in CL-bench requires models to learn from the context, ranging from new domain-specific knowledge, rule systems, and complex procedures to laws derived from empirical data, all of which are absent from pre-training. This goes far beyond long-context tasks that primarily test retrieval or reading comprehension, and in-context learning tasks, where models learn simple task patterns via instructions and demonstrations. Our evaluations of ten frontier LMs find that models solve only 17.2% of tasks on average. Even the best-performing model, GPT-5.1, solves only 23.7%, revealing that LMs have yet to achieve effective context learning, which poses a critical bottleneck for tackling real-world, complex context-dependent tasks. CL-bench represents a step towards building LMs with this fundamental capability, making them more intelligent and advancing their deployment in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03587
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CL-bench: A Benchmark for Context Learning
Dou, Shihan
Zhang, Ming
Yin, Zhangyue
Huang, Chenhao
Shen, Yujiong
Wang, Junzhe
Chen, Jiayi
Ni, Yuchen
Ye, Junjie
Zhang, Cheng
Xie, Huaibing
Hu, Jianglu
Wang, Shaolei
Wang, Weichao
Xiao, Yanling
Liu, Yiting
Xu, Zenan
Guo, Zhen
Zhou, Pluto
Gui, Tao
Wu, Zuxuan
Qiu, Xipeng
Zhang, Qi
Huang, Xuanjing
Jiang, Yu-Gang
Wang, Di
Yao, Shunyu
Computation and Language
Current language models (LMs) excel at reasoning over prompts using pre-trained knowledge. However, real-world tasks are far more complex and context-dependent: models must learn from task-specific context and leverage new knowledge beyond what is learned during pre-training to reason and resolve tasks. We term this capability context learning, a crucial ability that humans naturally possess but has been largely overlooked. To this end, we introduce CL-bench, a real-world benchmark consisting of 500 complex contexts, 1,899 tasks, and 31,607 verification rubrics, all crafted by experienced domain experts. Each task is designed such that the new content required to resolve it is contained within the corresponding context. Resolving tasks in CL-bench requires models to learn from the context, ranging from new domain-specific knowledge, rule systems, and complex procedures to laws derived from empirical data, all of which are absent from pre-training. This goes far beyond long-context tasks that primarily test retrieval or reading comprehension, and in-context learning tasks, where models learn simple task patterns via instructions and demonstrations. Our evaluations of ten frontier LMs find that models solve only 17.2% of tasks on average. Even the best-performing model, GPT-5.1, solves only 23.7%, revealing that LMs have yet to achieve effective context learning, which poses a critical bottleneck for tackling real-world, complex context-dependent tasks. CL-bench represents a step towards building LMs with this fundamental capability, making them more intelligent and advancing their deployment in real-world scenarios.
title CL-bench: A Benchmark for Context Learning
topic Computation and Language
url https://arxiv.org/abs/2602.03587