LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Yushi, Lv, Xin, Zhang, Jiajie, Lyu, Hongchang, Tang, Jiankai, Huang, Zhidian, Du, Zhengxiao, Liu, Xiao, Zeng, Aohan, Hou, Lei, Dong, Yuxiao, Tang, Jie, Li, Juanzi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
Understanding Emergent Abilities of Language Models from the Loss Perspective
by: Du, Zhengxiao, et al.
Published: (2024)
by: Du, Zhengxiao, et al.
Published: (2024)
LongAlign: A Recipe for Long Context Alignment of Large Language Models
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
by: Bai, Yushi, et al.
Published: (2026)
by: Bai, Yushi, et al.
Published: (2026)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
by: Zhong, Lucen, et al.
Published: (2025)
by: Zhong, Lucen, et al.
Published: (2025)
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
by: Zhang, Jiajie, et al.
Published: (2024)
by: Zhang, Jiajie, et al.
Published: (2024)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
by: Chen, Xueyao, et al.
Published: (2026)
by: Chen, Xueyao, et al.
Published: (2026)
LongReward: Improving Long-context Large Language Models with AI Feedback
by: Zhang, Jiajie, et al.
Published: (2024)
by: Zhang, Jiajie, et al.
Published: (2024)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
by: Wen, Haoming, et al.
Published: (2025)
by: Wen, Haoming, et al.
Published: (2025)
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
by: Lv, Minxuan, et al.
Published: (2026)
by: Lv, Minxuan, et al.
Published: (2026)
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
by: Sun, Kai, et al.
Published: (2025)
by: Sun, Kai, et al.
Published: (2025)
Pre-training Distillation for Large Language Models: A Design Space Exploration
by: Peng, Hao, et al.
Published: (2024)
by: Peng, Hao, et al.
Published: (2024)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
Shifting Long-Context LLMs Research from Input to Output
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2023)
by: Tu, Shangqing, et al.
Published: (2023)
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
by: Hou, Zhenyu, et al.
Published: (2025)
by: Hou, Zhenyu, et al.
Published: (2025)
LVBench: An Extreme Long Video Understanding Benchmark
by: Wang, Weihan, et al.
Published: (2024)
by: Wang, Weihan, et al.
Published: (2024)
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding
by: Jubair, Sheikh, et al.
Published: (2025)
by: Jubair, Sheikh, et al.
Published: (2025)
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
by: Qi, Ji, et al.
Published: (2024)
by: Qi, Ji, et al.
Published: (2024)
XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies
by: Ni, Xuanfan, et al.
Published: (2024)
by: Ni, Xuanfan, et al.
Published: (2024)
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
by: Huang, Zhongzhan, et al.
Published: (2025)
by: Huang, Zhongzhan, et al.
Published: (2025)
APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding
by: Liu, Mingdao, et al.
Published: (2024)
by: Liu, Mingdao, et al.
Published: (2024)
SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
by: Cui, Hao, et al.
Published: (2025)
by: Cui, Hao, et al.
Published: (2025)
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering
by: Zhao, Qingfei, et al.
Published: (2024)
by: Zhao, Qingfei, et al.
Published: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
MileBench: Benchmarking MLLMs in Long Context
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
by: Sun, Siqi, et al.
Published: (2026)
by: Sun, Siqi, et al.
Published: (2026)
LongSafety: Evaluating Long-Context Safety of Large Language Models
by: Lu, Yida, et al.
Published: (2025)
by: Lu, Yida, et al.
Published: (2025)
How do Transformers Learn Implicit Reasoning?
by: Ye, Jiaran, et al.
Published: (2025)
by: Ye, Jiaran, et al.
Published: (2025)
MuLD: The Multitask Long Document Benchmark
by: Hudson, G Thomas, et al.
Published: (2022)
by: Hudson, G Thomas, et al.
Published: (2022)
Similar Items
-
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
by: Bai, Yushi, et al.
Published: (2024) -
Understanding Emergent Abilities of Language Models from the Loss Perspective
by: Du, Zhengxiao, et al.
Published: (2024) -
LongAlign: A Recipe for Long Context Alignment of Large Language Models
by: Bai, Yushi, et al.
Published: (2024) -
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
by: Bai, Yushi, et al.
Published: (2024) -
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026)