How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yue, Wang, Yihao, Tang, Ziyi, Zheng, Yongsen, Wang, Keze |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChunQiuTR: Time-Keyed Temporal Retrieval in Classical Chinese Annals
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025)
by: Wang, Baode, et al.
Published: (2025)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025)
by: Wang, Baode, et al.
Published: (2025)
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Bridging Discourse Treebanks with a Unified Rhetorical Structure Parser
by: Chistova, Elena
Published: (2025)
by: Chistova, Elena
Published: (2025)
XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser
by: Cheng, Xianfu, et al.
Published: (2024)
by: Cheng, Xianfu, et al.
Published: (2024)
Self-Correction Makes LLMs Better Parsers
by: Zhang, Ziyan, et al.
Published: (2025)
by: Zhang, Ziyan, et al.
Published: (2025)
Llamipa: An Incremental Discourse Parser
by: Thompson, Kate, et al.
Published: (2024)
by: Thompson, Kate, et al.
Published: (2024)
SPAWNing Structural Priming Predictions from a Cognitively Motivated Parser
by: Prasad, Grusha, et al.
Published: (2024)
by: Prasad, Grusha, et al.
Published: (2024)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay
by: Tang, Ziyi, et al.
Published: (2025)
by: Tang, Ziyi, et al.
Published: (2025)
Document Structure in Long Document Transformers
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
by: Xiang, Biao, et al.
Published: (2026)
by: Xiang, Biao, et al.
Published: (2026)
Artificial Intelligence and the Spatial Documentation of Languages
by: Ghanim, Hakam
Published: (2024)
by: Ghanim, Hakam
Published: (2024)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)
by: Ma, Dongsheng, et al.
Published: (2026)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Do Multi-Document Summarization Models Synthesize?
by: DeYoung, Jay, et al.
Published: (2023)
by: DeYoung, Jay, et al.
Published: (2023)
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs
by: Wang, Zhizhi, et al.
Published: (2026)
by: Wang, Zhizhi, et al.
Published: (2026)
Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
HuAMR: A Hungarian AMR Parser and Dataset
by: Barta, Botond, et al.
Published: (2025)
by: Barta, Botond, et al.
Published: (2025)
How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning
by: Yuan, Jiahao, et al.
Published: (2026)
by: Yuan, Jiahao, et al.
Published: (2026)
Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey
by: Zhao, Yuqing, et al.
Published: (2026)
by: Zhao, Yuqing, et al.
Published: (2026)
DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation
by: Lyu, Xinglin, et al.
Published: (2025)
by: Lyu, Xinglin, et al.
Published: (2025)
Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
An Attempt to Develop a Neural Parser based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese
by: Nguyen, Duc-Vu, et al.
Published: (2024)
by: Nguyen, Duc-Vu, et al.
Published: (2024)
Clustering Document Parts: Detecting and Characterizing Influence Campaigns from Documents
by: Wang, Zhengxiang, et al.
Published: (2024)
by: Wang, Zhengxiang, et al.
Published: (2024)
Document Intelligence in the Era of Large Language Models: A Survey
by: Wang, Weishi, et al.
Published: (2025)
by: Wang, Weishi, et al.
Published: (2025)
A State-of-the-Art Morphosyntactic Parser and Lemmatizer for Ancient Greek
by: Celano, Giuseppe G. A.
Published: (2024)
by: Celano, Giuseppe G. A.
Published: (2024)
SETUP: Sentence-level English-To-Uniform Meaning Representation Parser
by: Markle, Emma, et al.
Published: (2025)
by: Markle, Emma, et al.
Published: (2025)
Syntactic Language Change in English and German: Metrics, Parsers, and Convergences
by: Chen, Yanran, et al.
Published: (2024)
by: Chen, Yanran, et al.
Published: (2024)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Single-Pass Document Scanning for Question Answering
by: Cao, Weili, et al.
Published: (2025)
by: Cao, Weili, et al.
Published: (2025)
READoc: A Unified Benchmark for Realistic Document Structured Extraction
by: Li, Zichao, et al.
Published: (2024)
by: Li, Zichao, et al.
Published: (2024)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens
by: Wang, Cunxiang, et al.
Published: (2024)
by: Wang, Cunxiang, et al.
Published: (2024)
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
by: Liang, Zhuowen, et al.
Published: (2026)
by: Liang, Zhuowen, et al.
Published: (2026)
ContextGuard: Structured Self-Auditing for Context Learning in Language Models
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
A Bionic Natural Language Parser Equivalent to a Pushdown Automaton
by: Wei, Zhenghao, et al.
Published: (2024)
by: Wei, Zhenghao, et al.
Published: (2024)
Document Reconstruction Unlocks Scalable Long-Context RLVR
by: Xiao, Yao, et al.
Published: (2026)
by: Xiao, Yao, et al.
Published: (2026)
Similar Items
-
ChunQiuTR: Time-Keyed Temporal Retrieval in Classical Chinese Annals
by: Wang, Yihao, et al.
Published: (2026) -
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025) -
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
by: Wang, Baode, et al.
Published: (2025) -
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
by: Wang, Yan, et al.
Published: (2025) -
DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
by: Zhang, Jusheng, et al.
Published: (2025)