SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Chelsi, Wu, Yiran, Zeng, Yifan, Liu, Jiale, Dai, S hengyu, Shao, Zhenwen, Wu, Qingyun, Wang, Huazheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory-Augmented Agent Training for Business Document Understanding
by: Liu, Jiale, et al.
Published: (2024)
by: Liu, Jiale, et al.
Published: (2024)
The Ranking Blind Spot: Decision Hijacking in LLM-based Text Ranking
by: Qian, Yaoyao, et al.
Published: (2025)
by: Qian, Yaoyao, et al.
Published: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scales
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
by: Zeng, Yifan, et al.
Published: (2026)
by: Zeng, Yifan, et al.
Published: (2026)
Live-Evo: Online Evolution of Agentic Memory from Continuous Feedback
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
by: Xiong, Junyu, et al.
Published: (2025)
by: Xiong, Junyu, et al.
Published: (2025)
LLM-RankFusion: Mitigating Intrinsic Inconsistency in LLM-based Ranking
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
by: Liu, Jiale, et al.
Published: (2025)
by: Liu, Jiale, et al.
Published: (2025)
Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search
by: Coalson, Zachary, et al.
Published: (2024)
by: Coalson, Zachary, et al.
Published: (2024)
PhyGaP: Physically-Grounded Gaussians with Polarization Cues
by: Wu, Jiale, et al.
Published: (2026)
by: Wu, Jiale, et al.
Published: (2026)
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
by: Ding, Yihao, et al.
Published: (2026)
by: Ding, Yihao, et al.
Published: (2026)
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025)
by: Rivas, Pablo, et al.
Published: (2025)
DBAutoDoc: Automated Discovery and Documentation of Undocumented Database Schemas via Statistical Analysis and Iterative LLM Refinement
by: Nagarajan, Amith, et al.
Published: (2026)
by: Nagarajan, Amith, et al.
Published: (2026)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
by: Zhen, Zongwei, et al.
Published: (2025)
by: Zhen, Zongwei, et al.
Published: (2025)
Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
by: Zhang, Shaokun, et al.
Published: (2025)
by: Zhang, Shaokun, et al.
Published: (2025)
MetaAgent-X : Breaking the Ceiling of Automatic Multi-Agent Systems via End-to-End Reinforcement Learning
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance Constraints
by: Xia, Xiaobo, et al.
Published: (2023)
by: Xia, Xiaobo, et al.
Published: (2023)
Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
by: Zheng, Yuanlei, et al.
Published: (2026)
by: Zheng, Yuanlei, et al.
Published: (2026)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
by: Qian, Kun, et al.
Published: (2025)
by: Qian, Kun, et al.
Published: (2025)
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
MemCollab: Cross-Model Memory Collaboration via Contrastive Trajectory Distillation
by: Chang, Yurui, et al.
Published: (2026)
by: Chang, Yurui, et al.
Published: (2026)
DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding
by: Wu, Yuchuan, et al.
Published: (2026)
by: Wu, Yuchuan, et al.
Published: (2026)
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
by: Bakkali, Souhail, et al.
Published: (2023)
by: Bakkali, Souhail, et al.
Published: (2023)
DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers
by: Wasserman, Navve, et al.
Published: (2025)
by: Wasserman, Navve, et al.
Published: (2025)
LLM-Augmented Retrieval: Enhancing Retrieval Models Through Language Models and Doc-Level Embedding
by: Wu, Mingrui, et al.
Published: (2024)
by: Wu, Mingrui, et al.
Published: (2024)
Adversarial Attacks on Combinatorial Multi-Armed Bandits
by: Balasubramanian, Rishab, et al.
Published: (2023)
by: Balasubramanian, Rishab, et al.
Published: (2023)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
by: Mo, Ye, et al.
Published: (2025)
by: Mo, Ye, et al.
Published: (2025)
DocReLM: Mastering Document Retrieval with Language Model
by: Wei, Gengchen, et al.
Published: (2024)
by: Wei, Gengchen, et al.
Published: (2024)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration
by: Tang, Yinghao, et al.
Published: (2026)
by: Tang, Yinghao, et al.
Published: (2026)
Embodied LLM Agents Learn to Cooperate in Organized Teams
by: Guo, Xudong, et al.
Published: (2024)
by: Guo, Xudong, et al.
Published: (2024)
A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
by: Yuan, Hui, et al.
Published: (2024)
by: Yuan, Hui, et al.
Published: (2024)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
by: Heakl, Ahmed, et al.
Published: (2026)
by: Heakl, Ahmed, et al.
Published: (2026)
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
by: Hu, Ruofan, et al.
Published: (2026)
by: Hu, Ruofan, et al.
Published: (2026)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
by: Wang, Wenjie, et al.
Published: (2026)
by: Wang, Wenjie, et al.
Published: (2026)
Randomized Iterative Solver as Iterative Refinement: A Simple Fix Towards Backward Stability
by: Xu, Ruihan, et al.
Published: (2024)
by: Xu, Ruihan, et al.
Published: (2024)
Similar Items
-
Memory-Augmented Agent Training for Business Document Understanding
by: Liu, Jiale, et al.
Published: (2024) -
The Ranking Blind Spot: Decision Hijacking in LLM-based Text Ranking
by: Qian, Yaoyao, et al.
Published: (2025) -
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024) -
EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scales
by: Zhang, Yaolun, et al.
Published: (2026) -
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
by: Zeng, Yifan, et al.
Published: (2026)