VERSE: Visual Embedding Reduction and Space Exploration. Clustering-Guided Insights for Training Data Enhancement in Visually-Rich Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | de Rodrigo, Ignacio, Lopez-Lopez, Alvaro J., Boal, Jaime |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents
by: Napolitano, Davide, et al.
Published: (2025)
by: Napolitano, Davide, et al.
Published: (2025)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
by: Li, Yueying, et al.
Published: (2026)
by: Li, Yueying, et al.
Published: (2026)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
by: Fujitake, Masato
Published: (2024)
by: Fujitake, Masato
Published: (2024)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
by: Lim, Sohwi, et al.
Published: (2026)
by: Lim, Sohwi, et al.
Published: (2026)
VERSE: Virtual-Gradient Aware Streaming Lifelong Learning with Anytime Inference
by: Banerjee, Soumya, et al.
Published: (2023)
by: Banerjee, Soumya, et al.
Published: (2023)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
by: Zhu, Wanrong, et al.
Published: (2024)
by: Zhu, Wanrong, et al.
Published: (2024)
Merging Context Clustering with Visual State Space Models for Medical Image Segmentation
by: Zhu, Yun, et al.
Published: (2025)
by: Zhu, Yun, et al.
Published: (2025)
Embedded Visual Prompt Tuning
by: Zu, Wenqiang, et al.
Published: (2024)
by: Zu, Wenqiang, et al.
Published: (2024)
On the Faithfulness of Visual Thinking: Measurement and Enhancement
by: Liu, Zujing, et al.
Published: (2025)
by: Liu, Zujing, et al.
Published: (2025)
Training-Free Text-Guided Image Editing with Visual Autoregressive Model
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
Relation-Rich Visual Document Generator for Visual Information Extraction
by: Jiang, Zi-Han, et al.
Published: (2025)
by: Jiang, Zi-Han, et al.
Published: (2025)
Powerful Teachers Matter: Text-Guided Multi-view Knowledge Distillation with Visual Prior Enhancement
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Towards Understanding Visual Grounding in Visual Language Models
by: Pantazopoulos, Georgios, et al.
Published: (2025)
by: Pantazopoulos, Georgios, et al.
Published: (2025)
Metric-Guided Feature Fusion of Visual Foundation Models for Segmentation Tasks
by: Guo, Yachan, et al.
Published: (2026)
by: Guo, Yachan, et al.
Published: (2026)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
by: Mo, Ye, et al.
Published: (2025)
by: Mo, Ye, et al.
Published: (2025)
Internalized Reasoning for Long-Context Visual Document Understanding
by: Veselka, Austin
Published: (2026)
by: Veselka, Austin
Published: (2026)
Visual Prompt Discovery via Semantic Exploration
by: Kim, Jaechang, et al.
Published: (2026)
by: Kim, Jaechang, et al.
Published: (2026)
InsightSee: Advancing Multi-agent Vision-Language Models for Enhanced Visual Understanding
by: Zhang, Huaxiang, et al.
Published: (2024)
by: Zhang, Huaxiang, et al.
Published: (2024)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
by: Chen, Ketong, et al.
Published: (2025)
by: Chen, Ketong, et al.
Published: (2025)
RAVE: Residual Vector Embedding for CLIP-Guided Backlit Image Enhancement
by: Gaintseva, Tatiana, et al.
Published: (2024)
by: Gaintseva, Tatiana, et al.
Published: (2024)
QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models
by: Chi, Tien-Yu, et al.
Published: (2025)
by: Chi, Tien-Yu, et al.
Published: (2025)
SEDMamba: Enhancing Selective State Space Modelling with Bottleneck Mechanism and Fine-to-Coarse Temporal Fusion for Efficient Error Detection in Robot-Assisted Surgery
by: Xu, Jialang, et al.
Published: (2024)
by: Xu, Jialang, et al.
Published: (2024)
OLIVE: Object Level In-Context Visual Embeddings
by: Ossowski, Timothy, et al.
Published: (2024)
by: Ossowski, Timothy, et al.
Published: (2024)
VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis
by: Symeonidis-Herzig, Alexandre, et al.
Published: (2025)
by: Symeonidis-Herzig, Alexandre, et al.
Published: (2025)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
by: Tien, Dong Nguyen, et al.
Published: (2025)
by: Tien, Dong Nguyen, et al.
Published: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
by: Tanaka, Ryota, et al.
Published: (2025)
by: Tanaka, Ryota, et al.
Published: (2025)
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
How to Train Your Long-Context Visual Document Model
by: Veselka, Austin
Published: (2026)
by: Veselka, Austin
Published: (2026)
Cluster Contrast for Unsupervised Visual Representation Learning
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Visual Space Optimization for Zero-shot Learning
by: Wang, Xinsheng, et al.
Published: (2019)
by: Wang, Xinsheng, et al.
Published: (2019)
Audio Visual Segmentation Through Text Embeddings
by: Lee, Kyungbok, et al.
Published: (2025)
by: Lee, Kyungbok, et al.
Published: (2025)
HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation
by: Gadi, Hari Krishna, et al.
Published: (2026)
by: Gadi, Hari Krishna, et al.
Published: (2026)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
by: Tartaglini, Alexa R., et al.
Published: (2025)
by: Tartaglini, Alexa R., et al.
Published: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Similar Items
-
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025) -
Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents
by: Napolitano, Davide, et al.
Published: (2025) -
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025) -
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024) -
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
by: Li, Yueying, et al.
Published: (2026)