CL-bench Life: Can Language Models Learn from Real-Life Context?
Fuente:
arXiv
Saved in:
| Main Authors: | Dou, Shihan, Shen, Yujiong, Huang, Chenhao, Ye, Junjie, Chen, Jiayi, Wang, Junzhe, He, Qianyu, Liu, Shichun, Lv, Changze, Lin, Jiahang, Zhang, Jiazheng, Zhang, Ming, Liu, Shaofan, Ji, Tao, Yin, Zhangyue, Zhang, Cheng, Xie, Huaibing, Hu, Jianglu, Deng, Jingcheng, Li, Lincheng, Hu, Minda, Wang, Shaolei, Zhao, Syrus, Wang, Weichao, Lei, Yan, Liu, Yang, Xiao, Yanling, Liu, Yiting, Xu, Zenan, Guo, Zhen, Zhao, Ziliang, Zhou, Pluto, Gui, Tao, Zhang, Qi, Huang, Xuanjing, Jiang, Yu-Gang, Wang, Di, Yao, Shunyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CL-bench: A Benchmark for Context Learning
by: Dou, Shihan, et al.
Published: (2026)
by: Dou, Shihan, et al.
Published: (2026)
Probing How Scalable Table Data Enhances General Long-Context Reasoning
by: Xie, Huaibing, et al.
Published: (2026)
by: Xie, Huaibing, et al.
Published: (2026)
A Decomposition Perspective to Long-context Reasoning for LLMs
by: Xiao, Yanling, et al.
Published: (2026)
by: Xiao, Yanling, et al.
Published: (2026)
EasyCraft: A Robust and Efficient Framework for Automatic Avatar Crafting
by: Wang, Suzhen, et al.
Published: (2025)
by: Wang, Suzhen, et al.
Published: (2025)
Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration
by: He, Bowei, et al.
Published: (2026)
by: He, Bowei, et al.
Published: (2026)
Empirical Analysis of Vulnerabilities Life Cycle in Golang Ecosystem
by: Hu, Jinchang, et al.
Published: (2023)
by: Hu, Jinchang, et al.
Published: (2023)
TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
by: Zhang, Ming, et al.
Published: (2024)
by: Zhang, Ming, et al.
Published: (2024)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
by: Pan, Chengjun, et al.
Published: (2026)
by: Pan, Chengjun, et al.
Published: (2026)
Organoids in gastrointestinal diseases: from bench to clinic
by: Qinying Wang, et al.
Published: (2024)
by: Qinying Wang, et al.
Published: (2024)
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
by: Zhang, Jiazheng, et al.
Published: (2026)
by: Zhang, Jiazheng, et al.
Published: (2026)
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
by: Zhang, Ming, et al.
Published: (2025)
by: Zhang, Ming, et al.
Published: (2025)
Recycling of Polydicyclopentadiene Enabled with N‐Coordinated Boronic Ester Bonds
by: Jiawei Hu, et al.
Published: (2024)
by: Jiawei Hu, et al.
Published: (2024)
Large Scale Unsupervised Brain MRI Image Registration Solution for Learn2Reg 2024
by: Zhang, Yuxi, et al.
Published: (2024)
by: Zhang, Yuxi, et al.
Published: (2024)
Incremental Structure Discovery of Classification via Sequential Monte Carlo
by: Huang, Changze, et al.
Published: (2024)
by: Huang, Changze, et al.
Published: (2024)
EgoLife: Towards Egocentric Life Assistant
by: Yang, Jingkang, et al.
Published: (2025)
by: Yang, Jingkang, et al.
Published: (2025)
SWE-bench Goes Live!
by: Zhang, Linghao, et al.
Published: (2025)
by: Zhang, Linghao, et al.
Published: (2025)
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
by: Zhang, Jingyi, et al.
Published: (2025)
by: Zhang, Jingyi, et al.
Published: (2025)
MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
by: Lin, Jiahang, et al.
Published: (2026)
by: Lin, Jiahang, et al.
Published: (2026)
SafeMove-RL: A Certifiable Reinforcement Learning Framework for Dynamic Motion Constraints in Trajectory Planning
by: Liu, Tengfei, et al.
Published: (2025)
by: Liu, Tengfei, et al.
Published: (2025)
Multiscale modeling for enhanced battery health analysis: Pathways to longevity
by: Kaiyi Yang, et al.
Published: (2024)
by: Kaiyi Yang, et al.
Published: (2024)
ICE: Interactive 3D Game Character Editing via Dialogue
by: Wu, Haoqian, et al.
Published: (2024)
by: Wu, Haoqian, et al.
Published: (2024)
Can RL Improve Generalization of LLM Agents? An Empirical Study
by: Xi, Zhiheng, et al.
Published: (2026)
by: Xi, Zhiheng, et al.
Published: (2026)
Scanning diffraction imaging without stable illumination and scan position information
by: Liu, Tao, et al.
Published: (2023)
by: Liu, Tao, et al.
Published: (2023)
Study on the Effects of Composite Catalysts on the Curing Process and Pot Life of the HTPB/HDI‐Trimer Binder System
by: Ma Hui, et al.
Published: (2025)
by: Ma Hui, et al.
Published: (2025)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
Fidelity-Imposed Displacement Editing for the Learn2Reg 2024 SHG-BF Challenge
by: Wang, Jiacheng, et al.
Published: (2024)
by: Wang, Jiacheng, et al.
Published: (2024)
EfficientDreamer: High-Fidelity and Robust 3D Creation via Orthogonal-view Diffusion Prior
by: Hu, Zhipeng, et al.
Published: (2023)
by: Hu, Zhipeng, et al.
Published: (2023)
Unsupervised Multimodal 3D Medical Image Registration with Multilevel Correlation Balanced Optimization
by: Wang, Jiazheng, et al.
Published: (2024)
by: Wang, Jiazheng, et al.
Published: (2024)
Optimal dismantling of directed networks
by: Liu, Xueming, et al.
Published: (2025)
by: Liu, Xueming, et al.
Published: (2025)
Branch Chain Variations Modulate Pyridine Derivative Adsorption for Long‐Life Zinc‐Ion Battery
by: Lei Xu, et al.
Published: (2025)
by: Lei Xu, et al.
Published: (2025)
Dual‐Anion‐Dominated Electrolyte Design Manipulating Coordination and Boron‐Rich Interphase for Self‐Healing and Long‐Life Mg Metal Anode
by: Lu Zhang, et al.
Published: (2025)
by: Lu Zhang, et al.
Published: (2025)
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
by: Zhang, Junjie, et al.
Published: (2025)
by: Zhang, Junjie, et al.
Published: (2025)
CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
by: Chen, Yuzhu, et al.
Published: (2025)
by: Chen, Yuzhu, et al.
Published: (2025)
Self-Demos: Eliciting Out-of-Demonstration Generalizability in Large Language Models
by: He, Wei, et al.
Published: (2024)
by: He, Wei, et al.
Published: (2024)
The Ontogenetic Development of Rapid Eye Movement Sleep in Early Life and Its Regulatory Mechanisms
by: Xiu‐Juan Zhang, et al.
Published: (2026)
by: Xiu‐Juan Zhang, et al.
Published: (2026)
Coupling Dead‐Lithium Reactivation and Interfacial Stabilization for Long‐Life Lithium Metal Batteries
by: Qiuxue Jian, et al.
Published: (2026)
by: Qiuxue Jian, et al.
Published: (2026)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
by: Wang, You, et al.
Published: (2025)
by: Wang, You, et al.
Published: (2025)
Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation
by: Lv, Changze, et al.
Published: (2026)
by: Lv, Changze, et al.
Published: (2026)
Protecting Zinc Electrodes with Glutarimide: A Breakthrough in Dendrite Prevention and Cycle Life Extension
by: Kuan Hu, et al.
Published: (2025)
by: Kuan Hu, et al.
Published: (2025)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
Similar Items
-
CL-bench: A Benchmark for Context Learning
by: Dou, Shihan, et al.
Published: (2026) -
Probing How Scalable Table Data Enhances General Long-Context Reasoning
by: Xie, Huaibing, et al.
Published: (2026) -
A Decomposition Perspective to Long-context Reasoning for LLMs
by: Xiao, Yanling, et al.
Published: (2026) -
EasyCraft: A Robust and Efficient Framework for Automatic Avatar Crafting
by: Wang, Suzhen, et al.
Published: (2025) -
Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration
by: He, Bowei, et al.
Published: (2026)