Vision Language Models and Document Understanding

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Killedar, Shreyash
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901775518269440
author Killedar, Shreyash
author_facet Killedar, Shreyash
contents <p><span><span>This research paper presents an evaluation framework for analyzing document understanding in Vision–Language Models using PDF documents. The proposed framework accepts a document in PDF format, preserves the semantic content, and systematically alters visual layout and formatting attributes such as text alignment, spacing, font styles, and structural organization. Vision–Language Models process these documents and generate responses for tasks including content comprehension, information extraction, and question answering. The framework integrates layout variation, content consistency, and response analysis to evaluate robustness and sensitivity across different document representations. Experimental evaluation demonstrates that model performance varies significantly with layout changes despite identical underlying content, indicating a dependence on visual structure in addition to semantic information</span></span></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18402832
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Vision Language Models and Document Understanding
Killedar, Shreyash
Document Analysis
vision language model
Multimodal Learning
pdf analysis
AI
Artificial intelligence
Artificial Intelligence/standards
Insemination, Artificial/classification
<p><span><span>This research paper presents an evaluation framework for analyzing document understanding in Vision–Language Models using PDF documents. The proposed framework accepts a document in PDF format, preserves the semantic content, and systematically alters visual layout and formatting attributes such as text alignment, spacing, font styles, and structural organization. Vision–Language Models process these documents and generate responses for tasks including content comprehension, information extraction, and question answering. The framework integrates layout variation, content consistency, and response analysis to evaluate robustness and sensitivity across different document representations. Experimental evaluation demonstrates that model performance varies significantly with layout changes despite identical underlying content, indicating a dependence on visual structure in addition to semantic information</span></span></p>
title Vision Language Models and Document Understanding
topic Document Analysis
vision language model
Multimodal Learning
pdf analysis
AI
Artificial intelligence
Artificial Intelligence/standards
Insemination, Artificial/classification
url https://doi.org/10.5281/zenodo.18402832