POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yuan, Zhao, Zhongyin, Tian, Le, Wang, Haicheng, Ye, Xubing, You, Yangxiu, Yu, Zilin, Wu, Chuhan, Zhou, Xiao, Yu, Yang, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026)
by: Zhao, Zhongyin, et al.
Published: (2026)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
by: Wang, Haicheng, et al.
Published: (2026)
by: Wang, Haicheng, et al.
Published: (2026)
POINTS: Improving Your Vision-language Model with Affordable Strategies
by: Liu, Yuan, et al.
Published: (2024)
by: Liu, Yuan, et al.
Published: (2024)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
by: Liu, Yuan, et al.
Published: (2024)
by: Liu, Yuan, et al.
Published: (2024)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
Rethinking Overlooked Aspects in Vision-Language Models
by: Liu, Yuan, et al.
Published: (2024)
by: Liu, Yuan, et al.
Published: (2024)
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
by: Ye, Xubing, et al.
Published: (2024)
by: Ye, Xubing, et al.
Published: (2024)
VoCo-LLaMA: Towards Vision Compression with Large Language Models
by: Ye, Xubing, et al.
Published: (2024)
by: Ye, Xubing, et al.
Published: (2024)
Source-Free Domain Adaptation with Vision-Language Prior
by: Tang, Song, et al.
Published: (2026)
by: Tang, Song, et al.
Published: (2026)
Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models
by: Zhou, Andy, et al.
Published: (2023)
by: Zhou, Andy, et al.
Published: (2023)
Vision-Language Model Selection and Reuse for Downstream Adaptation
by: Tan, Hao-Zhe, et al.
Published: (2025)
by: Tan, Hao-Zhe, et al.
Published: (2025)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
Distilling Vision-Language Models on Millions of Videos
by: Zhao, Yue, et al.
Published: (2024)
by: Zhao, Yue, et al.
Published: (2024)
Vision Harnessing Agent for Open Ad-hoc Segmentation
by: Wang, Zilin, et al.
Published: (2026)
by: Wang, Zilin, et al.
Published: (2026)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
by: Liu, Zhuchenyang, et al.
Published: (2026)
by: Liu, Zhuchenyang, et al.
Published: (2026)
Adversarial Prompt Distillation for Vision-Language Models
by: Luo, Lin, et al.
Published: (2024)
by: Luo, Lin, et al.
Published: (2024)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
by: Yuan, Qianhao, et al.
Published: (2026)
by: Yuan, Qianhao, et al.
Published: (2026)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Why better safe than sensitive
by: Haicheng Zhao
Published: (2024)
by: Haicheng Zhao
Published: (2024)
Free-Grained Hierarchical Visual Recognition
by: Park, Seulki, et al.
Published: (2025)
by: Park, Seulki, et al.
Published: (2025)
Bayesian Test-Time Adaptation for Vision-Language Models
by: Zhou, Lihua, et al.
Published: (2025)
by: Zhou, Lihua, et al.
Published: (2025)
User-Feedback-Driven Adaptation for Vision-and-Language Navigation
by: Yu, Yongqiang, et al.
Published: (2025)
by: Yu, Yongqiang, et al.
Published: (2025)
FAILURE POINTS IN THE PKI ARCHITECTURE
by: Radomir I. Prodanović
Published: (2017)
by: Radomir I. Prodanović
Published: (2017)
Accelerating Diffusion Decoders via Multi-Scale Sampling and One-Step Distillation
by: Wang, Chuhan, et al.
Published: (2026)
by: Wang, Chuhan, et al.
Published: (2026)
DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
by: Li, Haodong, et al.
Published: (2024)
by: Li, Haodong, et al.
Published: (2024)
SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
by: Chen, Hanqi, et al.
Published: (2025)
by: Chen, Hanqi, et al.
Published: (2025)
Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model
by: Bingtao, Zhou, et al.
Published: (2026)
by: Bingtao, Zhou, et al.
Published: (2026)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
by: Liu, Aiwei, et al.
Published: (2025)
by: Liu, Aiwei, et al.
Published: (2025)
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
by: Huang, Zilin, et al.
Published: (2024)
by: Huang, Zilin, et al.
Published: (2024)
Computation‐efficiency‐oriented resource optimization in NOMA‐assisted multi‐UAV‐enabled MEC networks
by: Qingde Liu, et al.
Published: (2024)
by: Qingde Liu, et al.
Published: (2024)
Domain Adaptation Through Task Distillation
by: Zhou, Brady, et al.
Published: (2020)
by: Zhou, Brady, et al.
Published: (2020)
NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
by: Li, Zilin, et al.
Published: (2025)
by: Li, Zilin, et al.
Published: (2025)
scReader: Prompting Large Language Models to Interpret scRNA-seq Data
by: Li, Cong, et al.
Published: (2024)
by: Li, Cong, et al.
Published: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation
by: Xu, Changhua, et al.
Published: (2026)
by: Xu, Changhua, et al.
Published: (2026)
Unsupervised Vision‐Based Structural Anomaly Detection and Localization with Reverse Knowledge Distillation
by: Xiaoming Lei, et al.
Published: (2024)
by: Xiaoming Lei, et al.
Published: (2024)
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
by: Zhang, Gongbo, et al.
Published: (2026)
by: Zhang, Gongbo, et al.
Published: (2026)
Distill-SODA: Distilling Self-Supervised Vision Transformer for Source-Free Open-Set Domain Adaptation in Computational Pathology
by: Vray, Guillaume, et al.
Published: (2023)
by: Vray, Guillaume, et al.
Published: (2023)
Training-Free Test-Time Adaptation with Brownian Distance Covariance in Vision-Language Models
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
Similar Items
-
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026) -
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
by: Wang, Haicheng, et al.
Published: (2026) -
POINTS: Improving Your Vision-language Model with Affordable Strategies
by: Liu, Yuan, et al.
Published: (2024) -
POINTS1.5: Building a Vision-Language Model towards Real World Applications
by: Liu, Yuan, et al.
Published: (2024) -
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026)