DocVLM: Make Your VLM an Efficient Reader
Fuente:
arXiv
Saved in:
| Main Authors: | Nacson, Mor Shpigel, Aberdam, Aviad, Ganz, Roy, Avraham, Elad Ben, Golts, Alona, Kittenplon, Yair, Mazor, Shai, Litman, Ron |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
by: Fhima, Jonathan, et al.
Published: (2024)
by: Fhima, Jonathan, et al.
Published: (2024)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
by: Lokesh, K, et al.
Published: (2026)
by: Lokesh, K, et al.
Published: (2026)
DREAM: Deep Research Evaluation with Agentic Metrics
by: Avraham, Elad Ben, et al.
Published: (2026)
by: Avraham, Elad Ben, et al.
Published: (2026)
The Implicit Bias of Gradient Descent on Separable Data
by: Soudry, Daniel, et al.
Published: (2017)
by: Soudry, Daniel, et al.
Published: (2017)
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
by: Buzaglo, Gon, et al.
Published: (2024)
by: Buzaglo, Gon, et al.
Published: (2024)
Text-to-Image Generation Via Energy-Based CLIP
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026)
by: Sharony, Elad, et al.
Published: (2026)
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation
by: Hsu, Benjamin, et al.
Published: (2024)
by: Hsu, Benjamin, et al.
Published: (2024)
DODO: Discrete OCR Diffusion Models
by: Man, Sean, et al.
Published: (2026)
by: Man, Sean, et al.
Published: (2026)
Enhancing Consistency-Based Image Generation via Adversarialy-Trained Classification and Energy-Based Discrimination
by: Golan, Shelly, et al.
Published: (2024)
by: Golan, Shelly, et al.
Published: (2024)
MyVLM: Personalizing VLMs for User-Specific Queries
by: Alaluf, Yuval, et al.
Published: (2024)
by: Alaluf, Yuval, et al.
Published: (2024)
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
by: Wu, Chengyue, et al.
Published: (2026)
by: Wu, Chengyue, et al.
Published: (2026)
VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving
by: Liu, Haichao, et al.
Published: (2025)
by: Liu, Haichao, et al.
Published: (2025)
PushupBench: Your VLM is not good at counting pushups
by: Li, Shengzhi, et al.
Published: (2026)
by: Li, Shengzhi, et al.
Published: (2026)
Similarity-Aware Token Pruning: Your VLM but Faster
by: Jeddi, Ahmadreza, et al.
Published: (2025)
by: Jeddi, Ahmadreza, et al.
Published: (2025)
VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
by: Abramovich, Ofir, et al.
Published: (2024)
by: Abramovich, Ofir, et al.
Published: (2024)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
by: Singh, Aditya Kumar, et al.
Published: (2026)
by: Singh, Aditya Kumar, et al.
Published: (2026)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
Entropy governs the structure and reactivity of water dissociation under electric fields
by: Litman, Yair, et al.
Published: (2025)
by: Litman, Yair, et al.
Published: (2025)
Re:Verse -- Can Your VLM Read a Manga?
by: Baranwal, Aaditya, et al.
Published: (2025)
by: Baranwal, Aaditya, et al.
Published: (2025)
Aligning Artificial Superintelligence via a Multi-Box Protocol
by: Negozio, Avraham Yair
Published: (2025)
by: Negozio, Avraham Yair
Published: (2025)
A Recipe for Improving Remote Sensing VLM Zero Shot Generalization
by: Barzilai, Aviad, et al.
Published: (2025)
by: Barzilai, Aviad, et al.
Published: (2025)
Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM
by: Yu, Lei, et al.
Published: (2025)
by: Yu, Lei, et al.
Published: (2025)
Multilingual VLM Training: Adapting an English-Trained VLM to French
by: Lahmi, Jules, et al.
Published: (2025)
by: Lahmi, Jules, et al.
Published: (2025)
CadVLM: Bridging Language and Vision in the Generation of Parametric CAD Sketches
by: Wu, Sifan, et al.
Published: (2024)
by: Wu, Sifan, et al.
Published: (2024)
Hybrid Decision Making via Conformal VLM-generated Guidance
by: Banerjee, Debodeep, et al.
Published: (2026)
by: Banerjee, Debodeep, et al.
Published: (2026)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
by: Xue, Xizhe, et al.
Published: (2024)
by: Xue, Xizhe, et al.
Published: (2024)
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
by: Seo, Minjae, et al.
Published: (2025)
by: Seo, Minjae, et al.
Published: (2025)
Upper bounds for the entropy in the cusp for one-parameter diagonal flows on $SL_{d}(\mathbb{R})/SL_{d}(\mathbb{Z})$
by: Mor, Ron
Published: (2023)
by: Mor, Ron
Published: (2023)
Bounding entropy for one-parameter diagonal flows on $SL_{d}(\mathbb{R})/SL_{d}(\mathbb{Z})$ using linear functionals
by: Mor, Ron
Published: (2023)
by: Mor, Ron
Published: (2023)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
by: Zhang, Zaiwei, et al.
Published: (2024)
by: Zhang, Zaiwei, et al.
Published: (2024)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
by: Chen, Tao, et al.
Published: (2025)
by: Chen, Tao, et al.
Published: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction
by: Khan, Muhammad Haris, et al.
Published: (2025)
by: Khan, Muhammad Haris, et al.
Published: (2025)
GoalVLM: VLM-driven Object Goal Navigation for Multi-Agent System
by: James, MoniJesu, et al.
Published: (2026)
by: James, MoniJesu, et al.
Published: (2026)
ORION: ORthonormal Text Encoding for Universal VLM AdaptatION
by: Chakraborty, Omprakash, et al.
Published: (2026)
by: Chakraborty, Omprakash, et al.
Published: (2026)
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
by: DatologyAI, et al.
Published: (2026)
by: DatologyAI, et al.
Published: (2026)
SALSA: Single-pass Autoregressive LLM Structured Classification
by: Berdichevsky, Ruslan, et al.
Published: (2025)
by: Berdichevsky, Ruslan, et al.
Published: (2025)
Similar Items
-
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024) -
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
by: Fhima, Jonathan, et al.
Published: (2024) -
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024) -
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
by: Lokesh, K, et al.
Published: (2026) -
DREAM: Deep Research Evaluation with Agentic Metrics
by: Avraham, Elad Ben, et al.
Published: (2026)