Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Junjie, Gu, Gefei, Zheng, Yanan, Yeung, Dit-Yan, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Long-context LLMs Struggle with Long In-context Learning
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
by: Wang, Lanrui, et al.
Published: (2025)
by: Wang, Lanrui, et al.
Published: (2025)
LongReward: Improving Long-context Large Language Models with AI Feedback
by: Zhang, Jiajie, et al.
Published: (2024)
by: Zhang, Jiajie, et al.
Published: (2024)
Calibrating Long-form Generations from Large Language Models
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
What is Wrong with Perplexity for Long-context Language Modeling?
by: Fang, Lizhe, et al.
Published: (2024)
by: Fang, Lizhe, et al.
Published: (2024)
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
by: Gavin, Shawn, et al.
Published: (2024)
by: Gavin, Shawn, et al.
Published: (2024)
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
by: Zhang, Jiajie, et al.
Published: (2024)
by: Zhang, Jiajie, et al.
Published: (2024)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
by: Ye, Xiao, et al.
Published: (2024)
by: Ye, Xiao, et al.
Published: (2024)
LongAttn: Selecting Long-context Training Data via Token-level Attention
by: Wu, Longyun, et al.
Published: (2025)
by: Wu, Longyun, et al.
Published: (2025)
Long-context Non-factoid Question Answering in Indic Languages
by: Mishra, Ritwik, et al.
Published: (2025)
by: Mishra, Ritwik, et al.
Published: (2025)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
by: Riddell, Martin, et al.
Published: (2024)
by: Riddell, Martin, et al.
Published: (2024)
SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Revisiting Long-context Modeling from Context Denoising Perspective
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
Large Language Models Can Self-Improve in Long-context Reasoning
by: Li, Siheng, et al.
Published: (2024)
by: Li, Siheng, et al.
Published: (2024)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
Hierarchical Document Refinement for Long-context Retrieval-augmented Generation
by: Jin, Jiajie, et al.
Published: (2025)
by: Jin, Jiajie, et al.
Published: (2025)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
by: Gao, Chaochen, et al.
Published: (2025)
by: Gao, Chaochen, et al.
Published: (2025)
Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model
by: Gao, Chaochen, et al.
Published: (2024)
by: Gao, Chaochen, et al.
Published: (2024)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
by: Ye, Xi, et al.
Published: (2025)
by: Ye, Xi, et al.
Published: (2025)
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
Curse of High Dimensionality Issue in Transformer for Long-context Modeling
by: Zhang, Shuhai, et al.
Published: (2025)
by: Zhang, Shuhai, et al.
Published: (2025)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
by: Shangguan, Ziyao, et al.
Published: (2024)
by: Shangguan, Ziyao, et al.
Published: (2024)
An Effective Framework to Help Large Language Models Handle Numeric-involved Long-context Tasks
by: Yu, Yijiong
Published: (2024)
by: Yu, Yijiong
Published: (2024)
Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
by: Yu, Yijiong, et al.
Published: (2024)
by: Yu, Yijiong, et al.
Published: (2024)
A Decomposition Perspective to Long-context Reasoning for LLMs
by: Xiao, Yanling, et al.
Published: (2026)
by: Xiao, Yanling, et al.
Published: (2026)
Long-context Reference-based MT Quality Estimation
by: Haq, Sami Ul, et al.
Published: (2025)
by: Haq, Sami Ul, et al.
Published: (2025)
ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty
by: Zhong, Meizhi, et al.
Published: (2024)
by: Zhong, Meizhi, et al.
Published: (2024)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
by: Ding, Jiayu, et al.
Published: (2025)
by: Ding, Jiayu, et al.
Published: (2025)
RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback
by: Tang, Qiaoyu, et al.
Published: (2025)
by: Tang, Qiaoyu, et al.
Published: (2025)
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
Similar Items
-
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024) -
Long-context LLMs Struggle with Long In-context Learning
by: Li, Tianle, et al.
Published: (2024) -
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024) -
NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
by: Wang, Lanrui, et al.
Published: (2025) -
LongReward: Improving Long-context Large Language Models with AI Feedback
by: Zhang, Jiajie, et al.
Published: (2024)