XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ni, Xuanfan, Cai, Hengyi, Wei, Xiaochi, Wang, Shuaiqiang, Yin, Dawei, Li, Piji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Systematic Evaluation of Large Language Models for Natural Language Generation Tasks
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024)
Towards Verifiable Text Generation with Evolving Memory and Self-Reflection
von: Sun, Hao, et al.
Veröffentlicht: (2023)
von: Sun, Hao, et al.
Veröffentlicht: (2023)
Tool Learning with Large Language Models: A Survey
von: Qu, Changle, et al.
Veröffentlicht: (2024)
von: Qu, Changle, et al.
Veröffentlicht: (2024)
AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions
von: Qu, Changle, et al.
Veröffentlicht: (2024)
von: Qu, Changle, et al.
Veröffentlicht: (2024)
Towards Completeness-Oriented Tool Retrieval for Large Language Models
von: Qu, Changle, et al.
Veröffentlicht: (2024)
von: Qu, Changle, et al.
Veröffentlicht: (2024)
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
von: Bai, Yushi, et al.
Veröffentlicht: (2023)
von: Bai, Yushi, et al.
Veröffentlicht: (2023)
Grounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation
von: Chen, Jiamin, et al.
Veröffentlicht: (2025)
von: Chen, Jiamin, et al.
Veröffentlicht: (2025)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
von: Wu, Yuhao, et al.
Veröffentlicht: (2024)
MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching
von: Qu, Changle, et al.
Veröffentlicht: (2026)
von: Qu, Changle, et al.
Veröffentlicht: (2026)
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
von: Huang, Zhongzhan, et al.
Veröffentlicht: (2025)
von: Huang, Zhongzhan, et al.
Veröffentlicht: (2025)
Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models
von: Liu, Wenhan, et al.
Veröffentlicht: (2024)
von: Liu, Wenhan, et al.
Veröffentlicht: (2024)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
Cross-model Control: Improving Multiple Large Language Models in One-time Training
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
von: Tian, Han, et al.
Veröffentlicht: (2026)
von: Tian, Han, et al.
Veröffentlicht: (2026)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
FltLM: An Intergrated Long-Context Large Language Model for Effective Context Filtering and Understanding
von: Deng, Jingyang, et al.
Veröffentlicht: (2024)
von: Deng, Jingyang, et al.
Veröffentlicht: (2024)
LLMs + Persona-Plug = Personalized LLMs
von: Liu, Jiongnan, et al.
Veröffentlicht: (2024)
von: Liu, Jiongnan, et al.
Veröffentlicht: (2024)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
LongEmbed: Extending Embedding Models for Long Context Retrieval
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
von: Ye, Xi, et al.
Veröffentlicht: (2025)
von: Ye, Xi, et al.
Veröffentlicht: (2025)
RRWKV: Capturing Long-range Dependencies in RWKV
von: Wang, Leilei
Veröffentlicht: (2023)
von: Wang, Leilei
Veröffentlicht: (2023)
LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams
von: Wu, Yongxuan, et al.
Veröffentlicht: (2025)
von: Wu, Yongxuan, et al.
Veröffentlicht: (2025)
Towards Threshold-Free KV Cache Pruning
von: Ni, Xuanfan, et al.
Veröffentlicht: (2025)
von: Ni, Xuanfan, et al.
Veröffentlicht: (2025)
TP-RAG: Benchmarking Retrieval-Augmented Large Language Model Agents for Spatiotemporal-Aware Travel Planning
von: Ni, Hang, et al.
Veröffentlicht: (2025)
von: Ni, Hang, et al.
Veröffentlicht: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
von: Dong, Jiancheng, et al.
Veröffentlicht: (2025)
von: Dong, Jiancheng, et al.
Veröffentlicht: (2025)
DBR: Divergence-Based Regularization for Debiasing Natural Language Understanding Models
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
von: Sun, Siqi, et al.
Veröffentlicht: (2026)
von: Sun, Siqi, et al.
Veröffentlicht: (2026)
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
von: Chen, Longze, et al.
Veröffentlicht: (2024)
von: Chen, Longze, et al.
Veröffentlicht: (2024)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
LooGLE: Can Long-Context Language Models Understand Long Contexts?
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
LongGenBench: Long-context Generation Benchmark
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
von: Liu, Xiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Systematic Evaluation of Large Language Models for Natural Language Generation Tasks
von: Ni, Xuanfan, et al.
Veröffentlicht: (2024) -
Towards Verifiable Text Generation with Evolving Memory and Self-Reflection
von: Sun, Hao, et al.
Veröffentlicht: (2023) -
Tool Learning with Large Language Models: A Survey
von: Qu, Changle, et al.
Veröffentlicht: (2024) -
AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning
von: Sun, Hao, et al.
Veröffentlicht: (2024) -
From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions
von: Qu, Changle, et al.
Veröffentlicht: (2024)