SagaScale: A Realistic, Scalable, and High-Quality Long-Context Benchmark Built from Full-Length Novels
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Guancheng, Hu, Yong, Wang, Wenqing, Yang, Yaming, Gao, Jiaheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States
by: Zhang, Zhenliang, et al.
Published: (2026)
by: Zhang, Zhenliang, et al.
Published: (2026)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024)
by: Liang, Jiashuo, et al.
Published: (2024)
Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
by: Hu, Xiang, et al.
Published: (2024)
by: Hu, Xiang, et al.
Published: (2024)
LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
by: Zhao, Liang, et al.
Published: (2024)
by: Zhao, Liang, et al.
Published: (2024)
A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts
by: Du, Wenzhuo, et al.
Published: (2025)
by: Du, Wenzhuo, et al.
Published: (2025)
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
Context Discipline and Performance Correlation: Analyzing LLM Performance and Quality Degradation Under Varying Context Lengths
by: Ponnusamy, Ahilan Ayyachamy Nadar, et al.
Published: (2025)
by: Ponnusamy, Ahilan Ayyachamy Nadar, et al.
Published: (2025)
NExtLong: Toward Effective Long-Context Training without Long Documents
by: Gao, Chaochen, et al.
Published: (2025)
by: Gao, Chaochen, et al.
Published: (2025)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
by: Li, Junlong, et al.
Published: (2025)
by: Li, Junlong, et al.
Published: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
by: Wang, Shuting, et al.
Published: (2025)
by: Wang, Shuting, et al.
Published: (2025)
MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
LongStory: Coherent, Complete and Length Controlled Long story Generation
by: Park, Kyeongman, et al.
Published: (2023)
by: Park, Kyeongman, et al.
Published: (2023)
WildLong: Synthesizing Realistic Long-Context Instruction Data at Scale
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
by: Yang, Penghui, et al.
Published: (2025)
by: Yang, Penghui, et al.
Published: (2025)
EntropyLong: Effective Long-Context Training via Predictive Uncertainty
by: Jia, Junlong, et al.
Published: (2025)
by: Jia, Junlong, et al.
Published: (2025)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
by: Nguyen, Huu, et al.
Published: (2025)
by: Nguyen, Huu, et al.
Published: (2025)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
by: He, Guangxin, et al.
Published: (2025)
by: He, Guangxin, et al.
Published: (2025)
Efficient Context Scaling with LongCat ZigZag Attention
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets
by: Joshi, Harshit, et al.
Published: (2026)
by: Joshi, Harshit, et al.
Published: (2026)
Beyond Single-Granularity Prompts: A Multi-Scale Chain-of-Thought Prompt Learning for Graph
by: Zheng, Ziyu, et al.
Published: (2025)
by: Zheng, Ziyu, et al.
Published: (2025)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
by: Du, Yufeng, et al.
Published: (2025)
by: Du, Yufeng, et al.
Published: (2025)
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing
by: Li, Kuan, et al.
Published: (2025)
by: Li, Kuan, et al.
Published: (2025)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
by: Sun, Siqi, et al.
Published: (2026)
by: Sun, Siqi, et al.
Published: (2026)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
by: Wang, Minzheng, et al.
Published: (2024)
by: Wang, Minzheng, et al.
Published: (2024)
AcademicEval: Live Long-Context LLM Benchmark
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
by: Lei, Fangyu, et al.
Published: (2025)
by: Lei, Fangyu, et al.
Published: (2025)
When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
by: Zheng, Yingming, et al.
Published: (2025)
by: Zheng, Yingming, et al.
Published: (2025)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels
by: Wei, Lingxiao, et al.
Published: (2024)
by: Wei, Lingxiao, et al.
Published: (2024)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
by: Lee, Changhun, et al.
Published: (2025)
by: Lee, Changhun, et al.
Published: (2025)
Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model
by: Gao, Chaochen, et al.
Published: (2024)
by: Gao, Chaochen, et al.
Published: (2024)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
Similar Items
-
SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States
by: Zhang, Zhenliang, et al.
Published: (2026) -
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026) -
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024) -
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024) -
Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
by: Hu, Xiang, et al.
Published: (2024)