DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Yongkun, Chen, Pinxuan, Ying, Xuye, Chen, Zhineng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
by: Bai, Weikang, et al.
Published: (2025)
by: Bai, Weikang, et al.
Published: (2025)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
by: Yang, Minglai, et al.
Published: (2026)
by: Yang, Minglai, et al.
Published: (2026)
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
by: Zhang, Yaping, et al.
Published: (2026)
by: Zhang, Yaping, et al.
Published: (2026)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
by: Ouyang, Linke, et al.
Published: (2024)
by: Ouyang, Linke, et al.
Published: (2024)
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model
by: Wu, Zhanglin, et al.
Published: (2025)
by: Wu, Zhanglin, et al.
Published: (2025)
STORM: End-to-End Referring Multi-Object Tracking in Videos
by: Lu, Zijia, et al.
Published: (2026)
by: Lu, Zijia, et al.
Published: (2026)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
ContourFormer: Real-Time Contour-Based End-to-End Instance Segmentation Transformer
by: Yao, Weiwei, et al.
Published: (2025)
by: Yao, Weiwei, et al.
Published: (2025)
iPad: Iterative Proposal-centric End-to-End Autonomous Driving
by: Guo, Ke, et al.
Published: (2025)
by: Guo, Ke, et al.
Published: (2025)
Adversarial Flow Matching for Imperceptible Attacks on End-to-End Autonomous Driving
by: Zeng, Xinyu, et al.
Published: (2026)
by: Zeng, Xinyu, et al.
Published: (2026)
Towards End-to-End Neuromorphic Event-based 3D Object Reconstruction Without Physical Priors
by: Xu, Chuanzhi, et al.
Published: (2025)
by: Xu, Chuanzhi, et al.
Published: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
ParkingScenes: A Structured Dataset for End-to-End Autonomous Parking in Simulation Scenes
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
by: Li, Zhang, et al.
Published: (2026)
by: Li, Zhang, et al.
Published: (2026)
End-to-End Human Instance Matting
by: Liu, Qinglin, et al.
Published: (2024)
by: Liu, Qinglin, et al.
Published: (2024)
USAD: End-to-End Human Activity Recognition via Diffusion Model with Spatiotemporal Attention
by: Xiao, Hang, et al.
Published: (2025)
by: Xiao, Hang, et al.
Published: (2025)
EA-RAS: Towards Efficient and Accurate End-to-End Reconstruction of Anatomical Skeleton
by: Peng, Zhiheng, et al.
Published: (2024)
by: Peng, Zhiheng, et al.
Published: (2024)
Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving
by: Ding, Kairui, et al.
Published: (2024)
by: Ding, Kairui, et al.
Published: (2024)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
by: Ye, Xingsong, et al.
Published: (2024)
by: Ye, Xingsong, et al.
Published: (2024)
Guiding Attention in End-to-End Driving Models
by: Porres, Diego, et al.
Published: (2024)
by: Porres, Diego, et al.
Published: (2024)
Without Paired Labeled Data: End-to-End Self-Supervised Learning for Drone-view Geo-Localization
by: Chen, Zhongwei, et al.
Published: (2025)
by: Chen, Zhongwei, et al.
Published: (2025)
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
by: Zheng, Chaoda, et al.
Published: (2026)
by: Zheng, Chaoda, et al.
Published: (2026)
DocShaDiffusion: Diffusion Model in Latent Space for Document Image Shadow Removal
by: Liu, Wenjie, et al.
Published: (2025)
by: Liu, Wenjie, et al.
Published: (2025)
CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning
by: Ang, Sining, et al.
Published: (2026)
by: Ang, Sining, et al.
Published: (2026)
TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment
by: Qin, Chunxia, et al.
Published: (2026)
by: Qin, Chunxia, et al.
Published: (2026)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
by: Mo, Ye, et al.
Published: (2025)
by: Mo, Ye, et al.
Published: (2025)
PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
by: Sun, Ting, et al.
Published: (2025)
by: Sun, Ting, et al.
Published: (2025)
DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training
by: Xin, Chen, et al.
Published: (2024)
by: Xin, Chen, et al.
Published: (2024)
Fose: Fusion of One-Step Diffusion and End-to-End Network for Pansharpening
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
An End-to-End, Segmentation-Free, Arabic Handwritten Recognition Model on KHATT
by: Aabed, Sondos, et al.
Published: (2024)
by: Aabed, Sondos, et al.
Published: (2024)
Beyond Hungarian: Match-Free Supervision for End-to-End Object Detection
by: Qiu, Shoumeng, et al.
Published: (2026)
by: Qiu, Shoumeng, et al.
Published: (2026)
Augmenting End-to-End Steering Angle Prediction with CAN Bus Data
by: Singh, Amit
Published: (2023)
by: Singh, Amit
Published: (2023)
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Decoder Pre-Training with only Text for Scene Text Recognition
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
by: Islam, Md Mofijul, et al.
Published: (2026)
by: Islam, Md Mofijul, et al.
Published: (2026)
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
by: Tang, Yihong, et al.
Published: (2025)
by: Tang, Yihong, et al.
Published: (2025)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
An End-to-End Deep Learning Generative Framework for Refinable Shape Matching and Generation
by: Kalaie, Soodeh, et al.
Published: (2024)
by: Kalaie, Soodeh, et al.
Published: (2024)
Similar Items
-
Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
by: Bai, Weikang, et al.
Published: (2025) -
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
by: Yang, Minglai, et al.
Published: (2026) -
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
by: Zhang, Yaping, et al.
Published: (2026) -
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
by: Ouyang, Linke, et al.
Published: (2024) -
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)