Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Yuanlei, Fu, Pei, Li, Hang, Wang, Ziyang, Zhang, Yuyi, Ruan, Wenyu, Zhang, Xiaojin, Wei, Zhongyu, Luo, Zhenbo, Luan, Jian, Chen, Wei, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
RSL-SQL: Robust Schema Linking in Text-to-SQL Generation
by: Cao, Zhenbiao, et al.
Published: (2024)
by: Cao, Zhenbiao, et al.
Published: (2024)
AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
by: Wei, Lai, et al.
Published: (2024)
by: Wei, Lai, et al.
Published: (2024)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
by: Qi, Yukun, et al.
Published: (2026)
by: Qi, Yukun, et al.
Published: (2026)
CauESC: A Causal Aware Model for Emotional Support Conversation
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
by: Nguyen, Khanh, et al.
Published: (2025)
by: Nguyen, Khanh, et al.
Published: (2025)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
by: Cui, Cheng, et al.
Published: (2026)
by: Cui, Cheng, et al.
Published: (2026)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
by: Xiong, Junyu, et al.
Published: (2025)
by: Xiong, Junyu, et al.
Published: (2025)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
by: Zhang, Yiru, et al.
Published: (2025)
by: Zhang, Yiru, et al.
Published: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
by: Kang, Hengrui, et al.
Published: (2025)
by: Kang, Hengrui, et al.
Published: (2025)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
by: Mo, Ye, et al.
Published: (2025)
by: Mo, Ye, et al.
Published: (2025)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation
by: Wu, Qinzhuo, et al.
Published: (2025)
by: Wu, Qinzhuo, et al.
Published: (2025)
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
by: Long, Zhuohan, et al.
Published: (2026)
by: Long, Zhuohan, et al.
Published: (2026)
Coarse-to-Fine Detection of Multiple Seams for Robotic Welding
by: Wei, Pengkun, et al.
Published: (2024)
by: Wei, Pengkun, et al.
Published: (2024)
DocDeshadower: Frequency-Aware Transformer for Document Shadow Removal
by: Zhou, Ziyang, et al.
Published: (2023)
by: Zhou, Ziyang, et al.
Published: (2023)
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration
by: Tang, Yinghao, et al.
Published: (2026)
by: Tang, Yinghao, et al.
Published: (2026)
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
by: Xu, Yinsong, et al.
Published: (2026)
by: Xu, Yinsong, et al.
Published: (2026)
Demonstrating ViviDoc: Generating Interactive Documents through Human-Agent Collaboration
by: Tang, Yinghao, et al.
Published: (2026)
by: Tang, Yinghao, et al.
Published: (2026)
Magical Touch: Transforming Raw Capacitive Streams into Expressive Hand-Touchscreen Interaction
by: Guo, Yuanlei, et al.
Published: (2026)
by: Guo, Yuanlei, et al.
Published: (2026)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Theoretical Analysis of Privacy Leakage in Trustworthy Federated Learning: A Perspective from Linear Algebra and Optimization Theory
by: Zhang, Xiaojin, et al.
Published: (2024)
by: Zhang, Xiaojin, et al.
Published: (2024)
Bridging Privacy and Robustness for Trustworthy Machine Learning
by: Zhang, Xiaojin, et al.
Published: (2024)
by: Zhang, Xiaojin, et al.
Published: (2024)
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
by: Jain, Chelsi, et al.
Published: (2025)
by: Jain, Chelsi, et al.
Published: (2025)
Beyond ESM2: Graph-Enhanced Protein Sequence Modeling with Efficient Clustering
by: Jiao, Shujian, et al.
Published: (2024)
by: Jiao, Shujian, et al.
Published: (2024)
Stepwise Informativeness Search for Efficient and Effective LLM Reasoning
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
Coarse-to-Fine Grounded Memory for LLM Agent Planning
by: Yang, Wei, et al.
Published: (2025)
by: Yang, Wei, et al.
Published: (2025)
MapLocNet: Coarse-to-Fine Feature Registration for Visual Re-Localization in Navigation Maps
by: Wu, Hang, et al.
Published: (2024)
by: Wu, Hang, et al.
Published: (2024)
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
by: Xie, Junfei, et al.
Published: (2026)
by: Xie, Junfei, et al.
Published: (2026)
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
by: Fang, Yiyang, et al.
Published: (2026)
by: Fang, Yiyang, et al.
Published: (2026)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Similar Items
-
AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale
by: Wang, Ziyang, et al.
Published: (2025) -
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026) -
RSL-SQL: Robust Schema Linking in Text-to-SQL Generation
by: Cao, Zhenbiao, et al.
Published: (2024) -
AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
by: Li, Haoxuan, et al.
Published: (2025) -
MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
by: Wei, Lai, et al.
Published: (2024)