GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Siqi, Shen, Yufan, Chen, Xiangnan, Chen, Jiayi, Ju, Hengwei, Duan, Haodong, Mao, Song, Zhou, Hongbin, Zhang, Bo, Fu, Bin, Cai, Pinlong, Wen, Licheng, Shi, Botian, Liu, Yong, Cai, Xinyu, Qiao, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images
von: Li, Siqi, et al.
Veröffentlicht: (2025)
von: Li, Siqi, et al.
Veröffentlicht: (2025)
RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
von: Fu, Daocheng, et al.
Veröffentlicht: (2025)
von: Fu, Daocheng, et al.
Veröffentlicht: (2025)
DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
von: Wen, Licheng, et al.
Veröffentlicht: (2023)
von: Wen, Licheng, et al.
Veröffentlicht: (2023)
KG-TRACES: Enhancing Large Language Models with Knowledge Graph-constrained Trajectory Reasoning and Attribution Supervision
von: Wu, Rong, et al.
Veröffentlicht: (2025)
von: Wu, Rong, et al.
Veröffentlicht: (2025)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
von: Liu, Junming, et al.
Veröffentlicht: (2025)
von: Liu, Junming, et al.
Veröffentlicht: (2025)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base
von: Mao, Song, et al.
Veröffentlicht: (2025)
von: Mao, Song, et al.
Veröffentlicht: (2025)
Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
von: Yang, Cheng, et al.
Veröffentlicht: (2025)
von: Yang, Cheng, et al.
Veröffentlicht: (2025)
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
von: Wu, Rong, et al.
Veröffentlicht: (2025)
von: Wu, Rong, et al.
Veröffentlicht: (2025)
HetaRAG: Hybrid Deep Retrieval-Augmented Generation across Heterogeneous Data Stores
von: Yan, Guohang, et al.
Veröffentlicht: (2025)
von: Yan, Guohang, et al.
Veröffentlicht: (2025)
From Ranking to Selection: A Simple but Efficient Dynamic Passage Selector for Retrieval Augmented Generation
von: Meng, Siyuan, et al.
Veröffentlicht: (2025)
von: Meng, Siyuan, et al.
Veröffentlicht: (2025)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2024)
O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
von: Mei, Jianbiao, et al.
Veröffentlicht: (2025)
von: Mei, Jianbiao, et al.
Veröffentlicht: (2025)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving
von: Mei, Jianbiao, et al.
Veröffentlicht: (2024)
von: Mei, Jianbiao, et al.
Veröffentlicht: (2024)
TrafficMCTS: A Closed-Loop Traffic Flow Generation Framework with Group-Based Monte Carlo Tree Search
von: Fu, Ze, et al.
Veröffentlicht: (2023)
von: Fu, Ze, et al.
Veröffentlicht: (2023)
OASim: an Open and Adaptive Simulator based on Neural Rendering for Autonomous Driving
von: Yan, Guohang, et al.
Veröffentlicht: (2024)
von: Yan, Guohang, et al.
Veröffentlicht: (2024)
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2025)
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2025)
VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models
von: Ren, Yufan, et al.
Veröffentlicht: (2025)
von: Ren, Yufan, et al.
Veröffentlicht: (2025)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
von: Li, Mo, et al.
Veröffentlicht: (2024)
von: Li, Mo, et al.
Veröffentlicht: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
LimSim Series: An Autonomous Driving Simulation Platform for Validation and Enhancement
von: Fu, Daocheng, et al.
Veröffentlicht: (2025)
von: Fu, Daocheng, et al.
Veröffentlicht: (2025)
LimSim++: A Closed-Loop Platform for Deploying Multimodal LLMs in Autonomous Driving
von: Fu, Daocheng, et al.
Veröffentlicht: (2024)
von: Fu, Daocheng, et al.
Veröffentlicht: (2024)
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
von: Li, Hongxiang, et al.
Veröffentlicht: (2025)
von: Li, Hongxiang, et al.
Veröffentlicht: (2025)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
von: Cai, Jie, et al.
Veröffentlicht: (2025)
von: Cai, Jie, et al.
Veröffentlicht: (2025)
Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations
von: Yang, Chengxu, et al.
Veröffentlicht: (2025)
von: Yang, Chengxu, et al.
Veröffentlicht: (2025)
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
SymDrive: Realistic and Controllable Driving Simulator via Symmetric Auto-regressive Online Restoration
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2025)
IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
von: Wang, Pengyu, et al.
Veröffentlicht: (2026)
von: Wang, Pengyu, et al.
Veröffentlicht: (2026)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
von: Kong, Fanqi, et al.
Veröffentlicht: (2025)
von: Kong, Fanqi, et al.
Veröffentlicht: (2025)
DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
von: Yang, Xuemeng, et al.
Veröffentlicht: (2024)
von: Yang, Xuemeng, et al.
Veröffentlicht: (2024)
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images
von: Li, Siqi, et al.
Veröffentlicht: (2025) -
RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
von: Fu, Daocheng, et al.
Veröffentlicht: (2025) -
DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
von: Wen, Licheng, et al.
Veröffentlicht: (2023) -
KG-TRACES: Enhancing Large Language Models with Knowledge Graph-constrained Trajectory Reasoning and Attribution Supervision
von: Wu, Rong, et al.
Veröffentlicht: (2025) -
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
von: Liu, Junming, et al.
Veröffentlicht: (2025)