DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yuchuan, Zhuo, Minghan, Fu, Teng, Zhao, Mengyang, Li, Bin, Xue, Xiangyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
von: Niu, Ke, et al.
Veröffentlicht: (2025)
von: Niu, Ke, et al.
Veröffentlicht: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
von: Fu, Teng, et al.
Veröffentlicht: (2025)
von: Fu, Teng, et al.
Veröffentlicht: (2025)
From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code Generation
von: Niu, Ke, et al.
Veröffentlicht: (2025)
von: Niu, Ke, et al.
Veröffentlicht: (2025)
IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
von: Zhao, Mengyang, et al.
Veröffentlicht: (2025)
von: Zhao, Mengyang, et al.
Veröffentlicht: (2025)
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
von: Fu, Teng, et al.
Veröffentlicht: (2025)
von: Fu, Teng, et al.
Veröffentlicht: (2025)
EAFormer: Scene Text Segmentation with Edge-Aware Transformers
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
von: Huang, Yibin, et al.
Veröffentlicht: (2025)
von: Huang, Yibin, et al.
Veröffentlicht: (2025)
ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language Models
von: Niu, Ke, et al.
Veröffentlicht: (2025)
von: Niu, Ke, et al.
Veröffentlicht: (2025)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
AMR-CCR: Anchored Modular Retrieval for Continual Chinese Character Recognition
von: Wu, Yuchuan, et al.
Veröffentlicht: (2026)
von: Wu, Yuchuan, et al.
Veröffentlicht: (2026)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
von: Chen, Ketong, et al.
Veröffentlicht: (2025)
von: Chen, Ketong, et al.
Veröffentlicht: (2025)
PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
Distribution Aligned Semantics Adaption for Lifelong Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
von: Mo, Ye, et al.
Veröffentlicht: (2025)
von: Mo, Ye, et al.
Veröffentlicht: (2025)
Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
von: Peng, Kaixin, et al.
Veröffentlicht: (2025)
von: Peng, Kaixin, et al.
Veröffentlicht: (2025)
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
von: Shi, Fan, et al.
Veröffentlicht: (2025)
von: Shi, Fan, et al.
Veröffentlicht: (2025)
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
Enhancing Video Inpainting with Aligned Frame Interval Guidance
von: Xie, Ming, et al.
Veröffentlicht: (2025)
von: Xie, Ming, et al.
Veröffentlicht: (2025)
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation
von: Tan, Jianwen, et al.
Veröffentlicht: (2025)
von: Tan, Jianwen, et al.
Veröffentlicht: (2025)
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
von: Jiang, Zhouqiang, et al.
Veröffentlicht: (2024)
von: Jiang, Zhouqiang, et al.
Veröffentlicht: (2024)
TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
von: Mao, Dongxing, et al.
Veröffentlicht: (2026)
von: Mao, Dongxing, et al.
Veröffentlicht: (2026)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
von: Jain, Chelsi, et al.
Veröffentlicht: (2025)
von: Jain, Chelsi, et al.
Veröffentlicht: (2025)
Zero-Shot Chinese Character Recognition with Hierarchical Multi-Granularity Image-Text Aligning
von: Zhu, Yinglian, et al.
Veröffentlicht: (2025)
von: Zhu, Yinglian, et al.
Veröffentlicht: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
von: Wu, Zhixuan, et al.
Veröffentlicht: (2026)
von: Wu, Zhixuan, et al.
Veröffentlicht: (2026)
Synthesizing Efficient Data with Diffusion Models for Person Re-Identification Pre-Training
von: Niu, Ke, et al.
Veröffentlicht: (2024)
von: Niu, Ke, et al.
Veröffentlicht: (2024)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
AIForge-Doc: A Benchmark for Detecting AI-Forged Tampering in Financial and Form Documents
von: Wu, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wu, Jiaqi, et al.
Veröffentlicht: (2026)
EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding
von: Zou, Kai, et al.
Veröffentlicht: (2026)
von: Zou, Kai, et al.
Veröffentlicht: (2026)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024) -
CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
von: Niu, Ke, et al.
Veröffentlicht: (2025) -
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
von: Kang, Hengrui, et al.
Veröffentlicht: (2025) -
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
von: Fu, Teng, et al.
Veröffentlicht: (2025) -
From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code Generation
von: Niu, Ke, et al.
Veröffentlicht: (2025)