LongGenBench: Long-context Generation Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Xiang, Dong, Peijie, Hu, Xuming, Chu, Xiaowen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
por: Wu, Yuhao, et al.
Publicado: (2024)
por: Wu, Yuhao, et al.
Publicado: (2024)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
por: Ye, Xiao, et al.
Publicado: (2024)
por: Ye, Xiao, et al.
Publicado: (2024)
Long-context LLMs Struggle with Long In-context Learning
por: Li, Tianle, et al.
Publicado: (2024)
por: Li, Tianle, et al.
Publicado: (2024)
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
por: Bai, Yushi, et al.
Publicado: (2024)
por: Bai, Yushi, et al.
Publicado: (2024)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
por: Chen, Ziyang, et al.
Publicado: (2026)
por: Chen, Ziyang, et al.
Publicado: (2026)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
por: Chen, Yuhao, et al.
Publicado: (2026)
por: Chen, Yuhao, et al.
Publicado: (2026)
LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
por: Gao, Chaochen, et al.
Publicado: (2025)
por: Gao, Chaochen, et al.
Publicado: (2025)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
por: Jiang, Ziyan, et al.
Publicado: (2024)
por: Jiang, Ziyan, et al.
Publicado: (2024)
MileBench: Benchmarking MLLMs in Long Context
por: Song, Dingjie, et al.
Publicado: (2024)
por: Song, Dingjie, et al.
Publicado: (2024)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
por: Deng, Chao, et al.
Publicado: (2024)
por: Deng, Chao, et al.
Publicado: (2024)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
por: Sun, Siqi, et al.
Publicado: (2026)
por: Sun, Siqi, et al.
Publicado: (2026)
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
por: Inc, Xiaohongshu
Publicado: (2026)
por: Inc, Xiaohongshu
Publicado: (2026)
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
por: Tang, Zhenheng, et al.
Publicado: (2025)
por: Tang, Zhenheng, et al.
Publicado: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
por: Yang, Wang, et al.
Publicado: (2025)
por: Yang, Wang, et al.
Publicado: (2025)
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
por: Shen, Yiting, et al.
Publicado: (2026)
por: Shen, Yiting, et al.
Publicado: (2026)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
por: Song, Yuanyi, et al.
Publicado: (2025)
por: Song, Yuanyi, et al.
Publicado: (2025)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
por: Wan, Luanbo, et al.
Publicado: (2025)
por: Wan, Luanbo, et al.
Publicado: (2025)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
por: He, Muyu, et al.
Publicado: (2026)
por: He, Muyu, et al.
Publicado: (2026)
Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
por: Liu, Kai, et al.
Publicado: (2025)
por: Liu, Kai, et al.
Publicado: (2025)
LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams
por: Wu, Yongxuan, et al.
Publicado: (2025)
por: Wu, Yongxuan, et al.
Publicado: (2025)
LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
por: Xiao, Zikai, et al.
Publicado: (2025)
por: Xiao, Zikai, et al.
Publicado: (2025)
AllMem: A Memory-centric Recipe for Efficient Long-context Modeling
por: Wang, Ziming, et al.
Publicado: (2026)
por: Wang, Ziming, et al.
Publicado: (2026)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
por: Tang, Zecheng, et al.
Publicado: (2026)
por: Tang, Zecheng, et al.
Publicado: (2026)
Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
por: Li, Junzhuo, et al.
Publicado: (2025)
por: Li, Junzhuo, et al.
Publicado: (2025)
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models
por: Thonet, Thibaut, et al.
Publicado: (2024)
por: Thonet, Thibaut, et al.
Publicado: (2024)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
por: Li, Shuyue Stella, et al.
Publicado: (2026)
por: Li, Shuyue Stella, et al.
Publicado: (2026)
Large Language Models Can Self-Improve in Long-context Reasoning
por: Li, Siheng, et al.
Publicado: (2024)
por: Li, Siheng, et al.
Publicado: (2024)
Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model
por: Gao, Chaochen, et al.
Publicado: (2024)
por: Gao, Chaochen, et al.
Publicado: (2024)
LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
por: Zhu, Liya, et al.
Publicado: (2025)
por: Zhu, Liya, et al.
Publicado: (2025)
Revisiting Long-context Modeling from Context Denoising Perspective
por: Tang, Zecheng, et al.
Publicado: (2025)
por: Tang, Zecheng, et al.
Publicado: (2025)
Extract, Match, and Score: An Evaluation Paradigm for Long Question-context-answer Triplets in Financial Analysis
por: Hu, Bo, et al.
Publicado: (2025)
por: Hu, Bo, et al.
Publicado: (2025)
In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided Search
por: Li, Huihan, et al.
Publicado: (2023)
por: Li, Huihan, et al.
Publicado: (2023)
A Decomposition Perspective to Long-context Reasoning for LLMs
por: Xiao, Yanling, et al.
Publicado: (2026)
por: Xiao, Yanling, et al.
Publicado: (2026)
Long Input Benchmark for Russian Analysis
por: Churin, Igor, et al.
Publicado: (2024)
por: Churin, Igor, et al.
Publicado: (2024)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
por: Ding, Jiayu, et al.
Publicado: (2025)
por: Ding, Jiayu, et al.
Publicado: (2025)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
por: Ren, Yiming, et al.
Publicado: (2026)
por: Ren, Yiming, et al.
Publicado: (2026)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
por: Lai, Kunfeng, et al.
Publicado: (2025)
por: Lai, Kunfeng, et al.
Publicado: (2025)
Generalizing From Short to Long: Effective Data Synthesis for Long-Context Instruction Tuning
por: Zhu, Wenhao, et al.
Publicado: (2025)
por: Zhu, Wenhao, et al.
Publicado: (2025)
Ejemplares similares
-
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
por: Wu, Yuhao, et al.
Publicado: (2024) -
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
por: Liu, Xiang, et al.
Publicado: (2025) -
AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
por: Ye, Xiao, et al.
Publicado: (2024) -
Long-context LLMs Struggle with Long In-context Learning
por: Li, Tianle, et al.
Publicado: (2024) -
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
por: Bai, Yushi, et al.
Publicado: (2024)