DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zou, Anni, Yu, Wenhao, Zhang, Hongming, Ma, Kaixin, Cai, Deng, Zhang, Zhuosheng, Zhao, Hai, Yu, Dong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LASER: LLM Agent with State-Space Exploration for Web Navigation
par: Ma, Kaixin, et autres
Publié: (2023)
par: Ma, Kaixin, et autres
Publié: (2023)
Generalizable Chain-of-Thought Prompting in Mixed-task Scenarios with Large Language Models
par: Zou, Anni, et autres
Publié: (2023)
par: Zou, Anni, et autres
Publié: (2023)
Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
par: Yu, Wenhao, et autres
Publié: (2023)
par: Yu, Wenhao, et autres
Publié: (2023)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
par: Fang, Tianqing, et autres
Publié: (2025)
par: Fang, Tianqing, et autres
Publié: (2025)
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
par: Lee, Hyunji, et autres
Publié: (2025)
par: Lee, Hyunji, et autres
Publié: (2025)
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
par: Zhang, Zhisong, et autres
Publié: (2025)
par: Zhang, Zhisong, et autres
Publié: (2025)
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
par: Ma, Xinbei, et autres
Publié: (2024)
par: Ma, Xinbei, et autres
Publié: (2024)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
par: Zhang, Ce, et autres
Publié: (2025)
par: Zhang, Ce, et autres
Publié: (2025)
GLaPE: Gold Label-agnostic Prompt Evaluation and Optimization for Large Language Model
par: Zhang, Xuanchang, et autres
Publié: (2024)
par: Zhang, Xuanchang, et autres
Publié: (2024)
Dense X Retrieval: What Retrieval Granularity Should We Use?
par: Chen, Tong, et autres
Publié: (2023)
par: Chen, Tong, et autres
Publié: (2023)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
par: Jia, Mengzhao, et autres
Publié: (2024)
par: Jia, Mengzhao, et autres
Publié: (2024)
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
par: He, Hongliang, et autres
Publié: (2024)
par: He, Hongliang, et autres
Publié: (2024)
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
par: He, Hongliang, et autres
Publié: (2024)
par: He, Hongliang, et autres
Publié: (2024)
Retrieval-augmented GUI Agents with Generative Guidelines
par: Xu, Ran, et autres
Publié: (2025)
par: Xu, Ran, et autres
Publié: (2025)
Towards End-to-End Open Conversational Machine Reading
par: Zhou, Sizhe, et autres
Publié: (2022)
par: Zhou, Sizhe, et autres
Publié: (2022)
DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
par: Jing, Liqiang, et autres
Publié: (2024)
par: Jing, Liqiang, et autres
Publié: (2024)
RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
par: Ouyang, Siru, et autres
Publié: (2024)
par: Ouyang, Siru, et autres
Publié: (2024)
Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
par: Ma, Xinbei, et autres
Publié: (2024)
par: Ma, Xinbei, et autres
Publié: (2024)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
par: Wang, Xiaoyang, et autres
Publié: (2025)
par: Wang, Xiaoyang, et autres
Publié: (2025)
Mitigating Misleading Chain-of-Thought Reasoning with Selective Filtering
par: Wu, Yexin, et autres
Publié: (2024)
par: Wu, Yexin, et autres
Publié: (2024)
BriLLM: Brain-inspired Large Language Model
par: Zhao, Hai, et autres
Publié: (2025)
par: Zhao, Hai, et autres
Publié: (2025)
Don't Throw Away Your Pretrained Model
par: Feng, Shangbin, et autres
Publié: (2025)
par: Feng, Shangbin, et autres
Publié: (2025)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
par: Diao, Lingxiao, et autres
Publié: (2025)
par: Diao, Lingxiao, et autres
Publié: (2025)
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
par: Tang, Xiangru, et autres
Publié: (2023)
par: Tang, Xiangru, et autres
Publié: (2023)
Self-Prompting Large Language Models for Zero-Shot Open-Domain QA
par: Li, Junlong, et autres
Publié: (2022)
par: Li, Junlong, et autres
Publié: (2022)
Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning
par: Zhao, Haodong, et autres
Publié: (2025)
par: Zhao, Haodong, et autres
Publié: (2025)
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
par: Guo, Yuan, et autres
Publié: (2025)
par: Guo, Yuan, et autres
Publié: (2025)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
par: Huang, Chengsong, et autres
Publié: (2025)
par: Huang, Chengsong, et autres
Publié: (2025)
Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
par: Zhang, Zhisong, et autres
Publié: (2024)
par: Zhang, Zhisong, et autres
Publié: (2024)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
par: Ma, Junyu, et autres
Publié: (2025)
par: Ma, Junyu, et autres
Publié: (2025)
MEGen: Generative Backdoor into Large Language Models via Model Editing
par: Qiu, Jiyang, et autres
Publié: (2024)
par: Qiu, Jiyang, et autres
Publié: (2024)
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
par: Wu, Di, et autres
Publié: (2024)
par: Wu, Di, et autres
Publié: (2024)
On the Robustness of Editing Large Language Models
par: Ma, Xinbei, et autres
Publié: (2024)
par: Ma, Xinbei, et autres
Publié: (2024)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
par: Yao, Yao, et autres
Publié: (2025)
par: Yao, Yao, et autres
Publié: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
par: Wang, Zhaowei, et autres
Publié: (2024)
par: Wang, Zhaowei, et autres
Publié: (2024)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
par: Yuan, Tongxin, et autres
Publié: (2024)
par: Yuan, Tongxin, et autres
Publié: (2024)
You Only Look at Screens: Multimodal Chain-of-Action Agents
par: Zhang, Zhuosheng, et autres
Publié: (2023)
par: Zhang, Zhuosheng, et autres
Publié: (2023)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
par: Shi, Yucheng, et autres
Publié: (2025)
par: Shi, Yucheng, et autres
Publié: (2025)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
par: Huang, Shulin, et autres
Publié: (2023)
par: Huang, Shulin, et autres
Publié: (2023)
Benchmark^2: Systematic Evaluation of LLM Benchmarks
par: Qian, Qi, et autres
Publié: (2026)
par: Qian, Qi, et autres
Publié: (2026)
Documents similaires
-
LASER: LLM Agent with State-Space Exploration for Web Navigation
par: Ma, Kaixin, et autres
Publié: (2023) -
Generalizable Chain-of-Thought Prompting in Mixed-task Scenarios with Large Language Models
par: Zou, Anni, et autres
Publié: (2023) -
Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
par: Yu, Wenhao, et autres
Publié: (2023) -
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
par: Fang, Tianqing, et autres
Publié: (2025) -
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
par: Lee, Hyunji, et autres
Publié: (2025)