DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhuang, Tianyi, Kuang, Chuqiao, Li, Xiaoguang, Teng, Yihua, Wu, Jihao, Wang, Yasheng, Shang, Lifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
von: Yan, Hao, et al.
Veröffentlicht: (2026)
von: Yan, Hao, et al.
Veröffentlicht: (2026)
SkeleGuide: Explicit Skeleton Reasoning for Context-Aware Human-in-Place Image Synthesis
von: Wu, Chuqiao, et al.
Veröffentlicht: (2026)
von: Wu, Chuqiao, et al.
Veröffentlicht: (2026)
PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
von: Tan, Haochen, et al.
Veröffentlicht: (2024)
von: Tan, Haochen, et al.
Veröffentlicht: (2024)
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
von: Tao, Xijia, et al.
Veröffentlicht: (2025)
von: Tao, Xijia, et al.
Veröffentlicht: (2025)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
DocFinQA: A Long-Context Financial Reasoning Dataset
von: Reddy, Varshini, et al.
Veröffentlicht: (2024)
von: Reddy, Varshini, et al.
Veröffentlicht: (2024)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning
von: Shi, Wenxuan, et al.
Veröffentlicht: (2025)
von: Shi, Wenxuan, et al.
Veröffentlicht: (2025)
Teaching Large Reasoning Models Effective Reflection
von: Wang, Hanbin, et al.
Veröffentlicht: (2026)
von: Wang, Hanbin, et al.
Veröffentlicht: (2026)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
von: Deng, Chao, et al.
Veröffentlicht: (2024)
von: Deng, Chao, et al.
Veröffentlicht: (2024)
Evaluating the External and Parametric Knowledge Fusion of Large Language Models
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
von: Gao, Fan, et al.
Veröffentlicht: (2025)
von: Gao, Fan, et al.
Veröffentlicht: (2025)
PuzzleJAX: A Benchmark for Reasoning and Learning
von: Earle, Sam, et al.
Veröffentlicht: (2025)
von: Earle, Sam, et al.
Veröffentlicht: (2025)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
von: Shankar, Shreya, et al.
Veröffentlicht: (2024)
von: Shankar, Shreya, et al.
Veröffentlicht: (2024)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning
von: Zhuang, Ren, et al.
Veröffentlicht: (2026)
von: Zhuang, Ren, et al.
Veröffentlicht: (2026)
DocSage: An Information Structuring Agent for Multi-Doc Multi-Entity Question Answering
von: Lin, Teng, et al.
Veröffentlicht: (2026)
von: Lin, Teng, et al.
Veröffentlicht: (2026)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
von: Qiu, Linlu, et al.
Veröffentlicht: (2023)
von: Qiu, Linlu, et al.
Veröffentlicht: (2023)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
von: Bishop, Jennifer A, et al.
Veröffentlicht: (2023)
von: Bishop, Jennifer A, et al.
Veröffentlicht: (2023)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
The Token Games: Evaluating Language Model Reasoning with Puzzle Duels
von: Henniger, Simon, et al.
Veröffentlicht: (2026)
von: Henniger, Simon, et al.
Veröffentlicht: (2026)
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
von: Zhang, Jiebin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2024)
The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
von: Ritchie, Logan, et al.
Veröffentlicht: (2026)
von: Ritchie, Logan, et al.
Veröffentlicht: (2026)
ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
von: Ying, Shuangshuang, et al.
Veröffentlicht: (2026)
von: Ying, Shuangshuang, et al.
Veröffentlicht: (2026)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
von: Liu, Chengwu, et al.
Veröffentlicht: (2025)
Tree of Agents: Improving Long-Context Capabilities of Large Language Models through Multi-Perspective Reasoning
von: Yu, Song, et al.
Veröffentlicht: (2025)
von: Yu, Song, et al.
Veröffentlicht: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking
von: Liu, Chang, et al.
Veröffentlicht: (2026) -
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
von: Yan, Hao, et al.
Veröffentlicht: (2026) -
SkeleGuide: Explicit Skeleton Reasoning for Context-Aware Human-in-Place Image Synthesis
von: Wu, Chuqiao, et al.
Veröffentlicht: (2026) -
PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
von: Tan, Haochen, et al.
Veröffentlicht: (2024) -
MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
von: Tao, Xijia, et al.
Veröffentlicht: (2025)