TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fhima, Jonathan, Avraham, Elad Ben, Nuriel, Oren, Kittenplon, Yair, Ganz, Roy, Aberdam, Aviad, Litman, Ron |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024)
by: Nacson, Mor Shpigel, et al.
Published: (2024)
DREAM: Deep Research Evaluation with Agentic Metrics
by: Avraham, Elad Ben, et al.
Published: (2026)
by: Avraham, Elad Ben, et al.
Published: (2026)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
Text-to-Image Generation Via Energy-Based CLIP
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation
by: Hsu, Benjamin, et al.
Published: (2024)
by: Hsu, Benjamin, et al.
Published: (2024)
Enhancing Consistency-Based Image Generation via Adversarialy-Trained Classification and Energy-Based Discrimination
by: Golan, Shelly, et al.
Published: (2024)
by: Golan, Shelly, et al.
Published: (2024)
Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Entropy governs the structure and reactivity of water dissociation under electric fields
by: Litman, Yair, et al.
Published: (2025)
by: Litman, Yair, et al.
Published: (2025)
Enhancing Retinal Vessel Segmentation Generalization via Layout-Aware Generative Modelling
by: Fhima, Jonathan, et al.
Published: (2025)
by: Fhima, Jonathan, et al.
Published: (2025)
Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
by: Bi, Tianci, et al.
Published: (2024)
by: Bi, Tianci, et al.
Published: (2024)
Aligning Artificial Superintelligence via a Multi-Box Protocol
by: Negozio, Avraham Yair
Published: (2025)
by: Negozio, Avraham Yair
Published: (2025)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
by: Mao, Zhiming, et al.
Published: (2024)
by: Mao, Zhiming, et al.
Published: (2024)
A Reality Check of Vision-Language Pre-training in Radiology: Have We Progressed Using Text?
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
by: Silva-Rodríguez, Julio, et al.
Published: (2025)
Diáspora y mestizaje en las novelas de Isaac Goldemberg
by: Patricia Nuriel
Published: (2008)
by: Patricia Nuriel
Published: (2008)
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
by: Zhong, Humen, et al.
Published: (2024)
by: Zhong, Humen, et al.
Published: (2024)
Paint by Inpaint: Learning to Add Image Objects by Removing Them First
by: Wasserman, Navve, et al.
Published: (2024)
by: Wasserman, Navve, et al.
Published: (2024)
Explaining Principles of Tip-Enhanced Raman Images with Ab Initio Modeling
by: Brezina, Krystof, et al.
Published: (2025)
by: Brezina, Krystof, et al.
Published: (2025)
PatchContrast: Self-Supervised Pre-training for 3D Object Detection
by: Shrout, Oren, et al.
Published: (2023)
by: Shrout, Oren, et al.
Published: (2023)
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
by: Šabanović, Ahmed, et al.
Published: (2026)
by: Šabanović, Ahmed, et al.
Published: (2026)
Identification of duloxetine to treat aniridia‐related keratopathy: Link to inflammation?
by: Irini Evnouchidou, et al.
Published: (2024)
by: Irini Evnouchidou, et al.
Published: (2024)
Identification of duloxetine to treat aniridia‐related keratopathy: Link to inflammation?
by: Irini Evnouchidou, et al.
Published: (2024)
by: Irini Evnouchidou, et al.
Published: (2024)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Ye, Wei, et al.
Published: (2024)
by: Ye, Wei, et al.
Published: (2024)
Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?
by: Che, Chengan, et al.
Published: (2026)
by: Che, Chengan, et al.
Published: (2026)
Class-Conditioned Transformation for Enhanced Robust Image Classification
by: Blau, Tsachi, et al.
Published: (2023)
by: Blau, Tsachi, et al.
Published: (2023)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
by: Luo, Gen, et al.
Published: (2024)
by: Luo, Gen, et al.
Published: (2024)
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
by: Jiang, Zhouqiang, et al.
Published: (2024)
by: Jiang, Zhouqiang, et al.
Published: (2024)
Development of a proptosis model as a surgical training tool for veterinary students and practitioners
by: Oren Pe'er, et al.
Published: (2025)
by: Oren Pe'er, et al.
Published: (2025)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
Unveiling the Deficiencies of Pre-trained Text-and-Layout Models in Real-world Visually-rich Document Information Extraction
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
SkinCLIP-VL: Consistency-Aware Vision-Language Learning for Multimodal Skin Cancer Diagnosis
by: Lu, Zhixiang, et al.
Published: (2026)
by: Lu, Zhixiang, et al.
Published: (2026)
Intention-Adaptive LLM Fine-Tuning for Text Revision Generation
by: Liu, Zhexiong, et al.
Published: (2026)
by: Liu, Zhexiong, et al.
Published: (2026)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
by: Li, Jianjian, et al.
Published: (2025)
by: Li, Jianjian, et al.
Published: (2025)
Decoupled Delay Compensation: Enhancing Pre-trained MARL Policies via Learned Dynamics Filtering
by: Mednikov, Maxim, et al.
Published: (2026)
by: Mednikov, Maxim, et al.
Published: (2026)
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
by: Liang, Mingliang, et al.
Published: (2026)
by: Liang, Mingliang, et al.
Published: (2026)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
Similar Items
-
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024) -
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024) -
DREAM: Deep Research Evaluation with Agentic Metrics
by: Avraham, Elad Ben, et al.
Published: (2026) -
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024) -
Text-to-Image Generation Via Energy-Based CLIP
by: Ganz, Roy, et al.
Published: (2024)