VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Abramovich, Ofir, Nayman, Niv, Fogel, Sharon, Lavi, Inbal, Litman, Ron, Tsiper, Shahar, Tichauer, Royee, Appalaraju, Srikar, Mazor, Shai, Manmatha, R. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DODO: Discrete OCR Diffusion Models
by: Man, Sean, et al.
Published: (2026)
by: Man, Sean, et al.
Published: (2026)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
FreeAugment: Data Augmentation Search Across All Degrees of Freedom
by: Bekor, Tom, et al.
Published: (2024)
by: Bekor, Tom, et al.
Published: (2024)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
Turbocharging Web Automation: The Impact of Compressed History States
by: Zhu, Xiyue, et al.
Published: (2025)
by: Zhu, Xiyue, et al.
Published: (2025)
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024)
by: Nacson, Mor Shpigel, et al.
Published: (2024)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Online Contract Design
by: Lavi, Elad, et al.
Published: (2026)
by: Lavi, Elad, et al.
Published: (2026)
Frequency-Aware Gaussian Splatting Decomposition
by: Lavi, Yishai, et al.
Published: (2025)
by: Lavi, Yishai, et al.
Published: (2025)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
by: Rao, Varun Nagaraj, et al.
Published: (2024)
by: Rao, Varun Nagaraj, et al.
Published: (2024)
Mocap Anywhere: Towards Pairwise-Distance based Motion Capture in the Wild (for the Wild)
by: Abramovich, Ofir, et al.
Published: (2026)
by: Abramovich, Ofir, et al.
Published: (2026)
Solving Convex Partition Visual Jigsaw Puzzles
by: Ohayon, Yaniv, et al.
Published: (2025)
by: Ohayon, Yaniv, et al.
Published: (2025)
Pictorial and apictorial polygonal jigsaw puzzles from arbitrary number of crossing cuts
by: Shahar, Peleg Harel Ofir Itzhak, et al.
Published: (2020)
by: Shahar, Peleg Harel Ofir Itzhak, et al.
Published: (2020)
The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragments
by: Shahar, Ofir Itzhak, et al.
Published: (2026)
by: Shahar, Ofir Itzhak, et al.
Published: (2026)
PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
by: Elkin, Gur, et al.
Published: (2025)
by: Elkin, Gur, et al.
Published: (2025)
Pairwise Alignment & Compatibility for Arbitrarily Irregular Image Fragments
by: Shahar, Ofir Itzhak, et al.
Published: (2025)
by: Shahar, Ofir Itzhak, et al.
Published: (2025)
Snake oil in action: Geographic and seasonal variability in epidermal lipids shape evaporative water loss in snakes
by: Shahar Dubiner, et al.
Published: (2025)
by: Shahar Dubiner, et al.
Published: (2025)
Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image Generation
by: Cohen, Nadav Z., et al.
Published: (2026)
by: Cohen, Nadav Z., et al.
Published: (2026)
Seasonal remodeling of visceral organs in the invasive desert gecko Tarentola annularis
by: Shahar DUBINER, et al.
Published: (2024)
by: Shahar DUBINER, et al.
Published: (2024)
An Overlooked Habitat‐Dependent Link Between Metabolism and Water Loss in Reptiles
by: Shahar Dubiner, et al.
Published: (2025)
by: Shahar Dubiner, et al.
Published: (2025)
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
by: Levy, Shahar, et al.
Published: (2025)
by: Levy, Shahar, et al.
Published: (2025)
EMAG: Differentiable 4D Gaussian Mixture Splatting for EEG Spatial Super-Resolution
by: Lazarovich, Alex, et al.
Published: (2026)
by: Lazarovich, Alex, et al.
Published: (2026)
A networked small-gain theorem based on discrete-time diagonal stability
by: Ofir, Ron, et al.
Published: (2024)
by: Ofir, Ron, et al.
Published: (2024)
Orthant-Monotonic Norms and Additive D-Stability
by: Ofir, Ron, et al.
Published: (2026)
by: Ofir, Ron, et al.
Published: (2026)
The $k$-Compound of a Difference-Algebraic System
by: Ofir, Ron, et al.
Published: (2021)
by: Ofir, Ron, et al.
Published: (2021)
Multiplicative and additive compounds via Kronecker products and Kronecker sums
by: Ofir, Ron, et al.
Published: (2024)
by: Ofir, Ron, et al.
Published: (2024)
Nonreciprocal surface plasmons in angularly varying, magnetized, metasurface tubes
by: Mazor, Yarden
Published: (2024)
by: Mazor, Yarden
Published: (2024)
CardioSpectrum: Comprehensive Myocardium Motion Analysis with 3D Deep Learning and Geometric Insights
by: Zuler, Shahar, et al.
Published: (2024)
by: Zuler, Shahar, et al.
Published: (2024)
Recognizing Artistic Style of Archaeological Image Fragments Using Deep Style Extrapolation
by: Elkin, Gur, et al.
Published: (2025)
by: Elkin, Gur, et al.
Published: (2025)
Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation
by: Abramovich, Ofir, et al.
Published: (2026)
by: Abramovich, Ofir, et al.
Published: (2026)
Automated Search for Conjectures on Mathematical Constants using Analysis of Integer Sequences
by: Razon, Ofir, et al.
Published: (2022)
by: Razon, Ofir, et al.
Published: (2022)
Tuning the Coherent Propagation of Organic Exciton-Polaritons through the Cavity Q-factor
by: Tichauer, Ruth H., et al.
Published: (2023)
by: Tichauer, Ruth H., et al.
Published: (2023)
Telehealth during the COVID‐19 pandemic: A positive hybrid model of therapeutic intervention in cerebral palsy
by: Lynne Fogel
Published: (2026)
by: Lynne Fogel
Published: (2026)
Paraguay: la constitución de la identidad femenina en el campo
by: Ramón Fogel
Published: (1994)
by: Ramón Fogel
Published: (1994)
La región de la triple frontera: territorios de integración y desintegración
by: Ramón Fogel
Published: (2008)
by: Ramón Fogel
Published: (2008)
Right-to-Act: A Pre-Execution Non-Compensatory Decision Protocol for AI Systems
by: Lavi, Gadi
Published: (2026)
by: Lavi, Gadi
Published: (2026)
One Person, One Bot
by: Lavi, Liat
Published: (2025)
by: Lavi, Liat
Published: (2025)
TextureSAM: Towards a Texture Aware Foundation Model for Segmentation
by: Cohen, Inbal, et al.
Published: (2025)
by: Cohen, Inbal, et al.
Published: (2025)
Similar Items
-
DODO: Discrete OCR Diffusion Models
by: Man, Sean, et al.
Published: (2026) -
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024) -
FreeAugment: Data Augmentation Search Across All Degrees of Freedom
by: Bekor, Tom, et al.
Published: (2024) -
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
by: Park, Joonhyung, et al.
Published: (2025) -
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024)