UPOCR: Towards Unified Pixel-Level OCR Interface
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Dezhi, Yang, Zhenhua, Zhang, Jiaxin, Liu, Chongyu, Shi, Yongxin, Ding, Kai, Guo, Fengjun, Jin, Lianwen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Predicting the Original Appearance of Damaged Historical Documents
by: Yang, Zhenhua, et al.
Published: (2024)
by: Yang, Zhenhua, et al.
Published: (2024)
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
by: Peng, Dezhi, et al.
Published: (2023)
by: Peng, Dezhi, et al.
Published: (2023)
Omni-IML: Towards Unified Image Manipulation Localization
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
by: Zhang, Yuyi, et al.
Published: (2025)
by: Zhang, Yuyi, et al.
Published: (2025)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Revisiting Tampered Scene Text Detection in the Era of Generative AI
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
MegaHan97K: A Large-Scale Dataset for Mega-Category Chinese Character Recognition with over 97K Categories
by: Zhang, Yuyi, et al.
Published: (2025)
by: Zhang, Yuyi, et al.
Published: (2025)
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
by: Zhang, Yuyi, et al.
Published: (2024)
by: Zhang, Yuyi, et al.
Published: (2024)
LEGO: Self-Supervised Representation Learning for Scene Text Images
by: Ren, Yujin, et al.
Published: (2024)
by: Ren, Yujin, et al.
Published: (2024)
Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography
by: Liu, Junle, et al.
Published: (2026)
by: Liu, Junle, et al.
Published: (2026)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
by: Liu, Yuliang, et al.
Published: (2023)
by: Liu, Yuliang, et al.
Published: (2023)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
by: Liu, Feiyang, et al.
Published: (2024)
by: Liu, Feiyang, et al.
Published: (2024)
Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
by: Zhang, Peirong, et al.
Published: (2024)
by: Zhang, Peirong, et al.
Published: (2024)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
by: Yang, Zhibo, et al.
Published: (2024)
by: Yang, Zhibo, et al.
Published: (2024)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
TextSleuth: Towards Explainable Tampered Text Detection
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
MCCD: A Multi-Attribute Chinese Calligraphy Character Dataset Annotated with Script Styles, Dynasties, and Calligraphers
by: Zhao, Yixin, et al.
Published: (2025)
by: Zhao, Yixin, et al.
Published: (2025)
Towards Pixel-Level VLM Perception via Simple Points Prediction
by: Song, Tianhui, et al.
Published: (2026)
by: Song, Tianhui, et al.
Published: (2026)
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
by: Pan, Wei, et al.
Published: (2025)
by: Pan, Wei, et al.
Published: (2025)
Agentar-Fin-OCR
by: Qian, Siyi, et al.
Published: (2026)
by: Qian, Siyi, et al.
Published: (2026)
Towards Pixel-Wise Anomaly Location for High-Resolution PCBA via Self-Supervised Image Reconstruction
by: Liu, Wuyi, et al.
Published: (2025)
by: Liu, Wuyi, et al.
Published: (2025)
An Empirical Study of Scaling Law for OCR
by: Rang, Miao, et al.
Published: (2023)
by: Rang, Miao, et al.
Published: (2023)
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
by: Huang, Xiaoshuang, et al.
Published: (2024)
by: Huang, Xiaoshuang, et al.
Published: (2024)
Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
by: Mao, Junyuan, et al.
Published: (2026)
by: Mao, Junyuan, et al.
Published: (2026)
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
by: Dong, Daxiang, et al.
Published: (2026)
by: Dong, Daxiang, et al.
Published: (2026)
Pixel is a Barrier: Diffusion Models Are More Adversarially Robust Than We Think
by: Xue, Haotian, et al.
Published: (2024)
by: Xue, Haotian, et al.
Published: (2024)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
by: Cai, Dexian, et al.
Published: (2025)
by: Cai, Dexian, et al.
Published: (2025)
JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Similar Items
-
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
by: Zhang, Jiaxin, et al.
Published: (2024) -
Predicting the Original Appearance of Damaged Historical Documents
by: Yang, Zhenhua, et al.
Published: (2024) -
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
by: Peng, Dezhi, et al.
Published: (2023) -
Omni-IML: Towards Unified Image Manipulation Localization
by: Qu, Chenfan, et al.
Published: (2024) -
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)