ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908530967052288 |
|---|---|
| author | Abdallah, Abdelrahman Mounis, Mohamed Abdalla, Mahmoud Kasem, Mahmoud SalahEldin Mahmoud, Mohamed Abdelhalim, Ibrahim Elkasaby, Mohamed ElBendary, Yasser Jatowt, Adam |
| author_facet | Abdallah, Abdelrahman Mounis, Mohamed Abdalla, Mahmoud Kasem, Mahmoud SalahEldin Mahmoud, Mohamed Abdelhalim, Ibrahim Elkasaby, Mohamed ElBendary, Yasser Jatowt, Adam |
| contents | Multilingual OCR and information extraction from receipts remains challenging, particularly for complex scripts like Arabic. We introduce \dataset, a comprehensive dataset designed for Arabic-English receipt understanding comprising 20,000 annotated receipts from diverse retail settings, 30,000 OCR-annotated images, and 10,000 item-level annotations, and a new Receipt QA subset with 1265 receipt images paired with 40 question-answer pairs each to support LLM evaluation for receipt understanding. The dataset captures merchant names, item descriptions, prices, receipt numbers, and dates to support object detection, OCR, and information extraction tasks. We establish baseline performance using traditional methods (Tesseract OCR) and advanced neural networks, demonstrating the dataset's effectiveness for processing complex, noisy real-world receipt layouts. Our publicly accessible dataset advances automated multilingual document processing research (see https://github.com/Update-For-Integrated-Business-AI/CORU ). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_04493 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding Abdallah, Abdelrahman Mounis, Mohamed Abdalla, Mahmoud Kasem, Mahmoud SalahEldin Mahmoud, Mohamed Abdelhalim, Ibrahim Elkasaby, Mohamed ElBendary, Yasser Jatowt, Adam Computer Vision and Pattern Recognition Computation and Language Multilingual OCR and information extraction from receipts remains challenging, particularly for complex scripts like Arabic. We introduce \dataset, a comprehensive dataset designed for Arabic-English receipt understanding comprising 20,000 annotated receipts from diverse retail settings, 30,000 OCR-annotated images, and 10,000 item-level annotations, and a new Receipt QA subset with 1265 receipt images paired with 40 question-answer pairs each to support LLM evaluation for receipt understanding. The dataset captures merchant names, item descriptions, prices, receipt numbers, and dates to support object detection, OCR, and information extraction tasks. We establish baseline performance using traditional methods (Tesseract OCR) and advanced neural networks, demonstrating the dataset's effectiveness for processing complex, noisy real-world receipt layouts. Our publicly accessible dataset advances automated multilingual document processing research (see https://github.com/Update-For-Integrated-Business-AI/CORU ). |
| title | ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2406.04493 |