CL-bench: A Benchmark for Context Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dou, Shihan, Zhang, Ming, Yin, Zhangyue, Huang, Chenhao, Shen, Yujiong, Wang, Junzhe, Chen, Jiayi, Ni, Yuchen, Ye, Junjie, Zhang, Cheng, Xie, Huaibing, Hu, Jianglu, Wang, Shaolei, Wang, Weichao, Xiao, Yanling, Liu, Yiting, Xu, Zenan, Guo, Zhen, Zhou, Pluto, Gui, Tao, Wu, Zuxuan, Qiu, Xipeng, Zhang, Qi, Huang, Xuanjing, Jiang, Yu-Gang, Wang, Di, Yao, Shunyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CL-bench Life: Can Language Models Learn from Real-Life Context?
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
Probing How Scalable Table Data Enhances General Long-Context Reasoning
von: Xie, Huaibing, et al.
Veröffentlicht: (2026)
von: Xie, Huaibing, et al.
Veröffentlicht: (2026)
A Decomposition Perspective to Long-context Reasoning for LLMs
von: Xiao, Yanling, et al.
Veröffentlicht: (2026)
von: Xiao, Yanling, et al.
Veröffentlicht: (2026)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
von: Lv, Huijie, et al.
Veröffentlicht: (2024)
von: Lv, Huijie, et al.
Veröffentlicht: (2024)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
Dynamic and Generalizable Process Reward Modeling
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training
von: Zhu, Dingwei, et al.
Veröffentlicht: (2026)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2026)
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
von: Wang, Weiyi, et al.
Veröffentlicht: (2026)
von: Wang, Weiyi, et al.
Veröffentlicht: (2026)
Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation
von: Lv, Changze, et al.
Veröffentlicht: (2026)
von: Lv, Changze, et al.
Veröffentlicht: (2026)
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
von: Zhiyuan, Zeng, et al.
Veröffentlicht: (2025)
von: Zhiyuan, Zeng, et al.
Veröffentlicht: (2025)
EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
von: Zang, Jianxiang, et al.
Veröffentlicht: (2025)
von: Zang, Jianxiang, et al.
Veröffentlicht: (2025)
Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
von: Yang, Yuming, et al.
Veröffentlicht: (2025)
von: Yang, Yuming, et al.
Veröffentlicht: (2025)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
Steering LLMs via Scalable Interactive Oversight
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
von: Liu, Boyang, et al.
Veröffentlicht: (2025)
von: Liu, Boyang, et al.
Veröffentlicht: (2025)
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
von: Li, Tianlong, et al.
Veröffentlicht: (2024)
von: Li, Tianlong, et al.
Veröffentlicht: (2024)
UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation
von: Li, Tianlong, et al.
Veröffentlicht: (2023)
von: Li, Tianlong, et al.
Veröffentlicht: (2023)
Domain Generalization via Causal Adjustment for Cross-Domain Sentiment Analysis
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions
von: Li, Peng, et al.
Veröffentlicht: (2026)
von: Li, Peng, et al.
Veröffentlicht: (2026)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models
von: Liu, Peiju, et al.
Veröffentlicht: (2026)
von: Liu, Peiju, et al.
Veröffentlicht: (2026)
Efficient Link Prediction via GNN Layers Induced by Negative Sampling
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CL-bench Life: Can Language Models Learn from Real-Life Context?
von: Dou, Shihan, et al.
Veröffentlicht: (2026) -
Probing How Scalable Table Data Enhances General Long-Context Reasoning
von: Xie, Huaibing, et al.
Veröffentlicht: (2026) -
A Decomposition Perspective to Long-context Reasoning for LLMs
von: Xiao, Yanling, et al.
Veröffentlicht: (2026) -
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
von: Sun, Yuhong, et al.
Veröffentlicht: (2025) -
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)