DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Renqiu, Mao, Song, Yan, Xiangchao, Zhou, Hongbin, Zhang, Bo, Peng, Haoyang, Pi, Jiahao, Fu, Daocheng, Wu, Wenjie, Ye, Hancheng, Feng, Shiyang, Wang, Bin, Xu, Chao, He, Conghui, Cai, Pinlong, Dou, Min, Shi, Botian, Zhou, Sheng, Wang, Yongwei, Yan, Junchi, Wu, Fei, Qiao, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
di: Ye, Hancheng, et al.
Pubblicazione: (2024)
di: Ye, Hancheng, et al.
Pubblicazione: (2024)
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
di: Xia, Renqiu, et al.
Pubblicazione: (2023)
di: Xia, Renqiu, et al.
Pubblicazione: (2023)
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
di: Yan, Xiangchao, et al.
Pubblicazione: (2025)
di: Yan, Xiangchao, et al.
Pubblicazione: (2025)
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
di: Fu, Daocheng, et al.
Pubblicazione: (2025)
di: Fu, Daocheng, et al.
Pubblicazione: (2025)
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
di: Yan, Xiangchao, et al.
Pubblicazione: (2023)
di: Yan, Xiangchao, et al.
Pubblicazione: (2023)
LimSim++: A Closed-Loop Platform for Deploying Multimodal LLMs in Autonomous Driving
di: Fu, Daocheng, et al.
Pubblicazione: (2024)
di: Fu, Daocheng, et al.
Pubblicazione: (2024)
KG-TRACES: Enhancing Large Language Models with Knowledge Graph-constrained Trajectory Reasoning and Attribution Supervision
di: Wu, Rong, et al.
Pubblicazione: (2025)
di: Wu, Rong, et al.
Pubblicazione: (2025)
DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
di: Wen, Licheng, et al.
Pubblicazione: (2023)
di: Wen, Licheng, et al.
Pubblicazione: (2023)
ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
di: Zhang, Bo, et al.
Pubblicazione: (2023)
di: Zhang, Bo, et al.
Pubblicazione: (2023)
AccidentGPT: Accident Analysis and Prevention from V2X Environmental Perception with Multi-modal Large Model
di: Wang, Lening, et al.
Pubblicazione: (2023)
di: Wang, Lening, et al.
Pubblicazione: (2023)
OASim: an Open and Adaptive Simulator based on Neural Rendering for Autonomous Driving
di: Yan, Guohang, et al.
Pubblicazione: (2024)
di: Yan, Guohang, et al.
Pubblicazione: (2024)
TrafficMCTS: A Closed-Loop Traffic Flow Generation Framework with Group-Based Monte Carlo Tree Search
di: Fu, Ze, et al.
Pubblicazione: (2023)
di: Fu, Ze, et al.
Pubblicazione: (2023)
SymDrive: Realistic and Controllable Driving Simulator via Symmetric Auto-regressive Online Restoration
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Liu, Zhiyuan, et al.
Pubblicazione: (2025)
Need for Mechanisms to Monitor Ocean Circulation‐Driven Seagrass Population Expansions
di: Zhaohua Wang, et al.
Pubblicazione: (2025)
di: Zhaohua Wang, et al.
Pubblicazione: (2025)
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2024)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
di: Ouyang, Linke, et al.
Pubblicazione: (2024)
di: Ouyang, Linke, et al.
Pubblicazione: (2024)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
di: Li, Zirui, et al.
Pubblicazione: (2025)
di: Li, Zirui, et al.
Pubblicazione: (2025)
LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval
di: Zhang, Yaoze, et al.
Pubblicazione: (2025)
di: Zhang, Yaoze, et al.
Pubblicazione: (2025)
Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
di: Wen, Zichen, et al.
Pubblicazione: (2025)
di: Wen, Zichen, et al.
Pubblicazione: (2025)
DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base
di: Mao, Song, et al.
Pubblicazione: (2025)
di: Mao, Song, et al.
Pubblicazione: (2025)
The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios
di: Fu, Daocheng, et al.
Pubblicazione: (2026)
di: Fu, Daocheng, et al.
Pubblicazione: (2026)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
di: Wang, Bin, et al.
Pubblicazione: (2024)
di: Wang, Bin, et al.
Pubblicazione: (2024)
Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward
di: Chen, Jianlong, et al.
Pubblicazione: (2026)
di: Chen, Jianlong, et al.
Pubblicazione: (2026)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
di: Wu, Rong, et al.
Pubblicazione: (2025)
di: Wu, Rong, et al.
Pubblicazione: (2025)
Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models
di: Huang, Jiaxi, et al.
Pubblicazione: (2025)
di: Huang, Jiaxi, et al.
Pubblicazione: (2025)
DocDancer: Towards Agentic Document-Grounded Information Seeking
di: Zhang, Qintong, et al.
Pubblicazione: (2026)
di: Zhang, Qintong, et al.
Pubblicazione: (2026)
LimSim Series: An Autonomous Driving Simulation Platform for Validation and Enhancement
di: Fu, Daocheng, et al.
Pubblicazione: (2025)
di: Fu, Daocheng, et al.
Pubblicazione: (2025)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
di: Yuan, Jiakang, et al.
Pubblicazione: (2025)
di: Yuan, Jiakang, et al.
Pubblicazione: (2025)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
di: Jia, Xiaosong, et al.
Pubblicazione: (2025)
di: Jia, Xiaosong, et al.
Pubblicazione: (2025)
On Reducing the Execution Latency of Superconducting Quantum Processors via Quantum Job Scheduling
di: Wu, Wenjie, et al.
Pubblicazione: (2024)
di: Wu, Wenjie, et al.
Pubblicazione: (2024)
Chimera: Improving Generalist Model with Domain-Specific Experts
di: Peng, Tianshuo, et al.
Pubblicazione: (2024)
di: Peng, Tianshuo, et al.
Pubblicazione: (2024)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
di: Kang, Hengrui, et al.
Pubblicazione: (2025)
di: Kang, Hengrui, et al.
Pubblicazione: (2025)
Dual Instruction Tuning with Large Language Models for Mathematical Reasoning
di: Zhou, Yongwei, et al.
Pubblicazione: (2024)
di: Zhou, Yongwei, et al.
Pubblicazione: (2024)
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
di: Liao, Wenhui, et al.
Pubblicazione: (2024)
di: Liao, Wenhui, et al.
Pubblicazione: (2024)
Quantum Circuit Synthesis and Compilation Optimization: Overview and Prospects
di: Yan, Ge, et al.
Pubblicazione: (2024)
di: Yan, Ge, et al.
Pubblicazione: (2024)
UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition
di: Wang, Bin, et al.
Pubblicazione: (2024)
di: Wang, Bin, et al.
Pubblicazione: (2024)
DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
di: Yang, Xuemeng, et al.
Pubblicazione: (2024)
di: Yang, Xuemeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
di: Xia, Renqiu, et al.
Pubblicazione: (2024) -
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
di: Ye, Hancheng, et al.
Pubblicazione: (2024) -
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
di: Xia, Renqiu, et al.
Pubblicazione: (2024) -
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
di: Xia, Renqiu, et al.
Pubblicazione: (2023) -
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
di: Yan, Xiangchao, et al.
Pubblicazione: (2025)