VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Humen, Yang, Zhibo, Li, Zhaohai, Wang, Peng, Tang, Jun, Cheng, Wenqing, Yao, Cong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Platypus: A Generalized Specialist Model for Reading Text in Various Forms
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
von: Yao, Yuan, et al.
Veröffentlicht: (2026)
von: Yao, Yuan, et al.
Veröffentlicht: (2026)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
Qwen2.5-VL Technical Report
von: Bai, Shuai, et al.
Veröffentlicht: (2025)
von: Bai, Shuai, et al.
Veröffentlicht: (2025)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
von: Wan, Jianqiang, et al.
Veröffentlicht: (2024)
von: Wan, Jianqiang, et al.
Veröffentlicht: (2024)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
Generative Compositor for Few-Shot Visual Information Extraction
von: Yang, Zhibo, et al.
Veröffentlicht: (2025)
von: Yang, Zhibo, et al.
Veröffentlicht: (2025)
Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
von: Zhao, Zhen, et al.
Veröffentlicht: (2023)
von: Zhao, Zhen, et al.
Veröffentlicht: (2023)
Qwen3-VL Technical Report
von: Bai, Shuai, et al.
Veröffentlicht: (2025)
von: Bai, Shuai, et al.
Veröffentlicht: (2025)
VL4AD: Vision-Language Models Improve Pixel-wise Anomaly Detection
von: Zhong, Liangyu, et al.
Veröffentlicht: (2024)
von: Zhong, Liangyu, et al.
Veröffentlicht: (2024)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
von: Li, Jianjian, et al.
Veröffentlicht: (2025)
von: Li, Jianjian, et al.
Veröffentlicht: (2025)
OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
von: Zeng, Lunbin, et al.
Veröffentlicht: (2025)
von: Zeng, Lunbin, et al.
Veröffentlicht: (2025)
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
von: Shen, Jinjie, et al.
Veröffentlicht: (2026)
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
von: Wei, Zhixiang, et al.
Veröffentlicht: (2026)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2026)
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
von: Dong, Daxiang, et al.
Veröffentlicht: (2025)
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
von: Hao, Zhiwei, et al.
Veröffentlicht: (2024)
von: Hao, Zhiwei, et al.
Veröffentlicht: (2024)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
von: Zhang, Yuyi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuyi, et al.
Veröffentlicht: (2024)
Visual Text Generation in the Wild
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
von: Zhang, Bo, et al.
Veröffentlicht: (2025)
von: Zhang, Bo, et al.
Veröffentlicht: (2025)
SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
von: Wang, Xiaokun, et al.
Veröffentlicht: (2025)
von: Wang, Xiaokun, et al.
Veröffentlicht: (2025)
An Effective Data Augmentation Method by Asking Questions about Scene Text Images
von: Yao, Xu, et al.
Veröffentlicht: (2026)
von: Yao, Xu, et al.
Veröffentlicht: (2026)
GraSP-VL: Length as a Semantic Granularity Interface for Vision-Language Representations
von: Li, Zesheng, et al.
Veröffentlicht: (2026)
von: Li, Zesheng, et al.
Veröffentlicht: (2026)
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
von: Thushara, Rusiru, et al.
Veröffentlicht: (2026)
von: Thushara, Rusiru, et al.
Veröffentlicht: (2026)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
von: Tang, Yutao, et al.
Veröffentlicht: (2025)
von: Tang, Yutao, et al.
Veröffentlicht: (2025)
DP^2-VL: Private Photo Dataset Protection by Data Poisoning for Vision-Language Models
von: Miao, Hongyi, et al.
Veröffentlicht: (2026)
von: Miao, Hongyi, et al.
Veröffentlicht: (2026)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
One Model for Two Tasks: Cooperatively Recognizing and Recovering Low-Resolution Scene Text Images by Iterative Mutual Guidance
von: Zhao, Minyi, et al.
Veröffentlicht: (2024)
von: Zhao, Minyi, et al.
Veröffentlicht: (2024)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Platypus: A Generalized Specialist Model for Reading Text in Various Forms
von: Wang, Peng, et al.
Veröffentlicht: (2024) -
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
von: Yao, Yuan, et al.
Veröffentlicht: (2026) -
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
von: Yang, Zhibo, et al.
Veröffentlicht: (2024) -
Qwen2.5-VL Technical Report
von: Bai, Shuai, et al.
Veröffentlicht: (2025) -
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)