Vision Language Models and Document Understanding
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Lingua: | inglese |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866901775518269440 |
|---|---|
| author | Killedar, Shreyash |
| author_facet | Killedar, Shreyash |
| contents | <p><span><span>This research paper presents an evaluation framework for analyzing document understanding in Vision–Language Models using PDF documents. The proposed framework accepts a document in PDF format, preserves the semantic content, and systematically alters visual layout and formatting attributes such as text alignment, spacing, font styles, and structural organization. Vision–Language Models process these documents and generate responses for tasks including content comprehension, information extraction, and question answering. The framework integrates layout variation, content consistency, and response analysis to evaluate robustness and sensitivity across different document representations. Experimental evaluation demonstrates that model performance varies significantly with layout changes despite identical underlying content, indicating a dependence on visual structure in addition to semantic information</span></span></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18402832 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Vision Language Models and Document Understanding Killedar, Shreyash Document Analysis vision language model Multimodal Learning pdf analysis AI Artificial intelligence Artificial Intelligence/standards Insemination, Artificial/classification <p><span><span>This research paper presents an evaluation framework for analyzing document understanding in Vision–Language Models using PDF documents. The proposed framework accepts a document in PDF format, preserves the semantic content, and systematically alters visual layout and formatting attributes such as text alignment, spacing, font styles, and structural organization. Vision–Language Models process these documents and generate responses for tasks including content comprehension, information extraction, and question answering. The framework integrates layout variation, content consistency, and response analysis to evaluate robustness and sensitivity across different document representations. Experimental evaluation demonstrates that model performance varies significantly with layout changes despite identical underlying content, indicating a dependence on visual structure in addition to semantic information</span></span></p> |
| title | Vision Language Models and Document Understanding |
| topic | Document Analysis vision language model Multimodal Learning pdf analysis AI Artificial intelligence Artificial Intelligence/standards Insemination, Artificial/classification |
| url | https://doi.org/10.5281/zenodo.18402832 |