TextBite: A Historical Czech Document Dataset for Logical Page Segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Kostelník, Martin, Beneš, Karel, Hradiš, Michal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
por: Kostelník, Martin, et al.
Publicado: (2026)
por: Kostelník, Martin, et al.
Publicado: (2026)
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
por: Kohút, Jan, et al.
Publicado: (2025)
por: Kohút, Jan, et al.
Publicado: (2025)
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
por: Kišš, Martin, et al.
Publicado: (2025)
por: Kišš, Martin, et al.
Publicado: (2025)
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
por: Kišš, Martin, et al.
Publicado: (2025)
por: Kišš, Martin, et al.
Publicado: (2025)
Practical Fine-Tuning of Autoregressive Models on Limited Handwritten Texts
por: Kohút, Jan, et al.
Publicado: (2025)
por: Kohút, Jan, et al.
Publicado: (2025)
Self-supervised Pre-training of Text Recognizers
por: Kišš, Martin, et al.
Publicado: (2024)
por: Kišš, Martin, et al.
Publicado: (2024)
Fine-tuning Is a Surprisingly Effective Domain Adaptation Baseline in Handwriting Recognition
por: Kohút, Jan, et al.
Publicado: (2023)
por: Kohút, Jan, et al.
Publicado: (2023)
Towards Writing Style Adaptation in Handwriting Recognition
por: Kohút, Jan, et al.
Publicado: (2023)
por: Kohút, Jan, et al.
Publicado: (2023)
Archival Faces: Detection of Faces in Digitized Historical Documents
por: Vaško, Marek, et al.
Publicado: (2025)
por: Vaško, Marek, et al.
Publicado: (2025)
PerBite: A Curated Diagnostic Workflow for Bite-Aware Food Volume Estimation
por: AlMughrabi, Ahmad, et al.
Publicado: (2026)
por: AlMughrabi, Ahmad, et al.
Publicado: (2026)
SegHist: A General Segmentation-based Framework for Chinese Historical Document Text Line Detection
por: Hu, Xingjian, et al.
Publicado: (2024)
por: Hu, Xingjian, et al.
Publicado: (2024)
Tuning-Free Amodal Segmentation via the Occlusion-Free Bias of Inpainting Models
por: Lee, Jae Joong, et al.
Publicado: (2025)
por: Lee, Jae Joong, et al.
Publicado: (2025)
CzechLynx: A Dataset for Individual Identification and Pose Estimation of the Eurasian Lynx
por: Picek, Lukas, et al.
Publicado: (2025)
por: Picek, Lukas, et al.
Publicado: (2025)
Few-Shot Connectivity-Aware Text Line Segmentation in Historical Documents
por: Sterzinger, Rafael, et al.
Publicado: (2025)
por: Sterzinger, Rafael, et al.
Publicado: (2025)
Top2Ground: A Height-Aware Dual Conditioning Diffusion Model for Robust Aerial-to-Ground View Generation
por: Lee, Jae Joong, et al.
Publicado: (2025)
por: Lee, Jae Joong, et al.
Publicado: (2025)
DELINE8K: A Synthetic Data Pipeline for the Semantic Segmentation of Historical Documents
por: Archibald, Taylor, et al.
Publicado: (2024)
por: Archibald, Taylor, et al.
Publicado: (2024)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
por: Xiong, Junyu, et al.
Publicado: (2025)
por: Xiong, Junyu, et al.
Publicado: (2025)
RGB2Point: 3D Point Cloud Generation from Single RGB Images
por: Lee, Jae Joong, et al.
Publicado: (2024)
por: Lee, Jae Joong, et al.
Publicado: (2024)
Accurate Fine-grained Layout Analysis for the Historical Tibetan Document Based on the Instance Segmentation
por: Zhao, Penghai, et al.
Publicado: (2021)
por: Zhao, Penghai, et al.
Publicado: (2021)
WAS: Dataset and Methods for Artistic Text Segmentation
por: Xie, Xudong, et al.
Publicado: (2024)
por: Xie, Xudong, et al.
Publicado: (2024)
Handheld Video Document Scanning: A Robust On-Device Model for Multi-Page Document Scanning
por: Wigington, Curtis
Publicado: (2024)
por: Wigington, Curtis
Publicado: (2024)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
por: Ortega, Marc Serra, et al.
Publicado: (2025)
por: Ortega, Marc Serra, et al.
Publicado: (2025)
Historical Printed Ornaments: Dataset and Tasks
por: Chaki, Sayan Kumar, et al.
Publicado: (2024)
por: Chaki, Sayan Kumar, et al.
Publicado: (2024)
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
por: Kasuba, Badri Vishal, et al.
Publicado: (2025)
por: Kasuba, Badri Vishal, et al.
Publicado: (2025)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
por: Li, Zongmin, et al.
Publicado: (2026)
por: Li, Zongmin, et al.
Publicado: (2026)
μgat: Improving Single-Page Document Parsing by Providing Multi-Page Context
por: Quattrini, Fabio, et al.
Publicado: (2024)
por: Quattrini, Fabio, et al.
Publicado: (2024)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
por: Do, Thao, et al.
Publicado: (2024)
por: Do, Thao, et al.
Publicado: (2024)
Tree-D Fusion: Simulation-Ready Tree Dataset from Single Images with Diffusion Priors
por: Lee, Jae Joong, et al.
Publicado: (2024)
por: Lee, Jae Joong, et al.
Publicado: (2024)
Handwriting Recognition in Historical Documents with Multimodal LLM
por: Li, Lucian
Publicado: (2024)
por: Li, Lucian
Publicado: (2024)
Predicting the Original Appearance of Damaged Historical Documents
por: Yang, Zhenhua, et al.
Publicado: (2024)
por: Yang, Zhenhua, et al.
Publicado: (2024)
LIGHT: Multi-Modal Text Linking on Historical Maps
por: Lin, Yijun, et al.
Publicado: (2025)
por: Lin, Yijun, et al.
Publicado: (2025)
A Dataset for the Recognition of Historical and Handwritten Music Scores in Western Notation
por: Torras, Pau, et al.
Publicado: (2026)
por: Torras, Pau, et al.
Publicado: (2026)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
por: Kang, Lei, et al.
Publicado: (2024)
por: Kang, Lei, et al.
Publicado: (2024)
Unsupervised Class Generation to Expand Semantic Segmentation Datasets
por: Montalvo, Javier, et al.
Publicado: (2025)
por: Montalvo, Javier, et al.
Publicado: (2025)
Improving OCR for Historical Texts of Multiple Languages
por: Westerdijk, Hylke, et al.
Publicado: (2025)
por: Westerdijk, Hylke, et al.
Publicado: (2025)
Hyper-Local Deformable Transformers for Text Spotting on Historical Maps
por: Lin, Yijun, et al.
Publicado: (2025)
por: Lin, Yijun, et al.
Publicado: (2025)
SynthmanticLiDAR: A Synthetic Dataset for Semantic Segmentation on LiDAR Imaging
por: Montalvo, Javier, et al.
Publicado: (2025)
por: Montalvo, Javier, et al.
Publicado: (2025)
CSAD: Unsupervised Component Segmentation for Logical Anomaly Detection
por: Hsieh, Yu-Hsuan, et al.
Publicado: (2024)
por: Hsieh, Yu-Hsuan, et al.
Publicado: (2024)
Sesame Plant Segmentation Dataset: A YOLO Formatted Annotated Dataset
por: Muhammad, Sunusi Ibrahim, et al.
Publicado: (2026)
por: Muhammad, Sunusi Ibrahim, et al.
Publicado: (2026)
Automatic Uncertainty-Aware Synthetic Data Bootstrapping for Historical Map Segmentation
por: Arzoumanidis, Lukas, et al.
Publicado: (2025)
por: Arzoumanidis, Lukas, et al.
Publicado: (2025)
Ejemplares similares
-
CzechTopic: A Benchmark for Zero-Shot Topic Localization in Historical Czech Documents
por: Kostelník, Martin, et al.
Publicado: (2026) -
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
por: Kohút, Jan, et al.
Publicado: (2025) -
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
por: Kišš, Martin, et al.
Publicado: (2025) -
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
por: Kišš, Martin, et al.
Publicado: (2025) -
Practical Fine-Tuning of Autoregressive Models on Limited Handwritten Texts
por: Kohút, Jan, et al.
Publicado: (2025)