Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
Fuente:
arXiv
Saved in:
| Main Authors: | Mansour, Adnan Ben, Karine, Ayoub, Naccache, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The future of document indexing: GPT and Donut revolutionize table of content processing
by: Feyisa, Degaga Wolde, et al.
Published: (2024)
by: Feyisa, Degaga Wolde, et al.
Published: (2024)
Robust feature knowledge distillation for enhanced performance of lightweight crack segmentation models
by: Chen, Zhaohui, et al.
Published: (2024)
by: Chen, Zhaohui, et al.
Published: (2024)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)
by: Zhang, Nonghai, et al.
Published: (2025)
I2CKD : Intra- and Inter-Class Knowledge Distillation for Semantic Segmentation
by: Karine, Ayoub, et al.
Published: (2024)
by: Karine, Ayoub, et al.
Published: (2024)
VLMs Guided Interpretable Decision Making for Autonomous Driving
by: Hu, Xin, et al.
Published: (2025)
by: Hu, Xin, et al.
Published: (2025)
MedConcept: Unsupervised Concept Discovery for Interpretability in Medical VLMs
by: Haque, Md Rakibul, et al.
Published: (2026)
by: Haque, Md Rakibul, et al.
Published: (2026)
Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru
by: Cusipuma, Dunant, et al.
Published: (2025)
by: Cusipuma, Dunant, et al.
Published: (2025)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
Learning effective pruning at initialization from iterative pruning
by: Liu, Shengkai, et al.
Published: (2024)
by: Liu, Shengkai, et al.
Published: (2024)
Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models
by: Bazi, Yakoub, et al.
Published: (2026)
by: Bazi, Yakoub, et al.
Published: (2026)
RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation
by: Zhang, Sen, et al.
Published: (2026)
by: Zhang, Sen, et al.
Published: (2026)
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
Interpreting Biomedical VLMs on High-Imbalance Out-of-Distributions: An Insight into BiomedCLIP on Radiology
by: Sadman, Nafiz, et al.
Published: (2025)
by: Sadman, Nafiz, et al.
Published: (2025)
Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
by: Xie, Zhuofan, et al.
Published: (2026)
by: Xie, Zhuofan, et al.
Published: (2026)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
Trustworthy Few-Shot Transfer of Medical VLMs through Split Conformal Prediction
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
by: Chen, Yixiong, et al.
Published: (2026)
by: Chen, Yixiong, et al.
Published: (2026)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
by: Wang, Xingyue, et al.
Published: (2026)
by: Wang, Xingyue, et al.
Published: (2026)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
by: Aich, Abhishek, et al.
Published: (2026)
by: Aich, Abhishek, et al.
Published: (2026)
LeCoT: revisiting network architecture for two-view correspondence pruning
by: Dai, Luanyuan, et al.
Published: (2025)
by: Dai, Luanyuan, et al.
Published: (2025)
Effective pruning of web-scale datasets based on complexity of concept clusters
by: Abbas, Amro, et al.
Published: (2024)
by: Abbas, Amro, et al.
Published: (2024)
Positioning radiata pine branches requiring pruning by drone stereo vision
by: Lin, Yida, et al.
Published: (2026)
by: Lin, Yida, et al.
Published: (2026)
Supporting Vision-Language Model Inference with Confounder-pruning Knowledge Prompt
by: Li, Jiangmeng, et al.
Published: (2022)
by: Li, Jiangmeng, et al.
Published: (2022)
Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
by: Liu, Yibo, et al.
Published: (2024)
by: Liu, Yibo, et al.
Published: (2024)
Exploring OCR-augmented Generation for Bilingual VQA
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Knowledge Condensation and Reasoning for Knowledge-based VQA
by: Hao, Dongze, et al.
Published: (2024)
by: Hao, Dongze, et al.
Published: (2024)
Measuring Faithful and Plausible Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2023)
by: Reich, Daniel, et al.
Published: (2023)
Deep Pre-Alignment for VLMs
by: Yu, Tianyu, et al.
Published: (2026)
by: Yu, Tianyu, et al.
Published: (2026)
Rapidly deploying on-device eye tracking by distilling visual foundation models
by: Jiang, Cheng, et al.
Published: (2026)
by: Jiang, Cheng, et al.
Published: (2026)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024)
by: Kahl, Kim-Celine, et al.
Published: (2024)
Progressive trajectory matching for medical dataset distillation
by: Yu, Zhen, et al.
Published: (2024)
by: Yu, Zhen, et al.
Published: (2024)
Attention to detail: inter-resolution knowledge distillation
by: del Amor, Rocío, et al.
Published: (2024)
by: del Amor, Rocío, et al.
Published: (2024)
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
by: Byun, Ji Young, et al.
Published: (2026)
by: Byun, Ji Young, et al.
Published: (2026)
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Neural network relief: a pruning algorithm based on neural activity
by: Dekhovich, Aleksandr, et al.
Published: (2021)
by: Dekhovich, Aleksandr, et al.
Published: (2021)
HLTCOE Evaluation Team at TREC 2025: VQA Track
by: Zhang, Dengjia, et al.
Published: (2025)
by: Zhang, Dengjia, et al.
Published: (2025)
SplatTalk: 3D VQA with Gaussian Splatting
by: Thai, Anh, et al.
Published: (2025)
by: Thai, Anh, et al.
Published: (2025)
Similar Items
-
The future of document indexing: GPT and Donut revolutionize table of content processing
by: Feyisa, Degaga Wolde, et al.
Published: (2024) -
Robust feature knowledge distillation for enhanced performance of lightweight crack segmentation models
by: Chen, Zhaohui, et al.
Published: (2024) -
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025) -
I2CKD : Intra- and Inter-Class Knowledge Distillation for Semantic Segmentation
by: Karine, Ayoub, et al.
Published: (2024) -
VLMs Guided Interpretable Decision Making for Autonomous Driving
by: Hu, Xin, et al.
Published: (2025)