Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Vivoli, Emanuele, Campaioli, Irene, Nardoni, Mariateresa, Biondi, Niccolò, Bertini, Marco, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
One missing piece in Vision and Language: A Survey on Comics Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
ComicsPAP: understanding comic strips by picking the correct panel
by: Vivoli, Emanuele, et al.
Published: (2025)
by: Vivoli, Emanuele, et al.
Published: (2025)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025)
by: Ortega, Marc Serra, et al.
Published: (2025)
HoloMine: A Synthetic Dataset for Buried Landmines Recognition using Microwave Holographic Imaging
by: Vivoli, Emanuele, et al.
Published: (2025)
by: Vivoli, Emanuele, et al.
Published: (2025)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
by: Gomez, Lluis, et al.
Published: (2014)
by: Gomez, Lluis, et al.
Published: (2014)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025)
by: Mishra, Pritam, et al.
Published: (2025)
Retrieval Augmented Comic Image Generation
by: Shui, Yunhao, et al.
Published: (2025)
by: Shui, Yunhao, et al.
Published: (2025)
Unlocking Comics: The AI4VA Dataset for Visual Understanding
by: Grönquist, Peter, et al.
Published: (2024)
by: Grönquist, Peter, et al.
Published: (2024)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
by: Chen, Yule, et al.
Published: (2025)
by: Chen, Yule, et al.
Published: (2025)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
by: Li, Yingxuan, et al.
Published: (2023)
by: Li, Yingxuan, et al.
Published: (2023)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025)
by: Sachdeva, Ragav, et al.
Published: (2025)
Towards Faithful Reasoning in Comics for Small MLLMs
by: Feng, Chengcheng, et al.
Published: (2026)
by: Feng, Chengcheng, et al.
Published: (2026)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
Predicting Winning Captions for Weekly New Yorker Comics
by: Cao, Stanley, et al.
Published: (2024)
by: Cao, Stanley, et al.
Published: (2024)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
by: Li, Zongmin, et al.
Published: (2026)
by: Li, Zongmin, et al.
Published: (2026)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
by: Li, Yingxuan, et al.
Published: (2024)
by: Li, Yingxuan, et al.
Published: (2024)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
by: Ryan, Yuriel, et al.
Published: (2025)
by: Ryan, Yuriel, et al.
Published: (2025)
FRED: The Florence RGB-Event Drone Dataset
by: Magrini, Gabriele, et al.
Published: (2025)
by: Magrini, Gabriele, et al.
Published: (2025)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
Labeling Comic Mischief Content in Online Videos with a Multimodal Hierarchical-Cross-Attention Model
by: Baharlouei, Elaheh, et al.
Published: (2024)
by: Baharlouei, Elaheh, et al.
Published: (2024)
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026)
by: Fu, Xuanshuo, et al.
Published: (2026)
Reading Between the Lanes: Text VideoQA on the Road
by: Tom, George, et al.
Published: (2023)
by: Tom, George, et al.
Published: (2023)
DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models
by: Kwon, Patrick, et al.
Published: (2025)
by: Kwon, Patrick, et al.
Published: (2025)
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024)
by: Llabrés, Artemis, et al.
Published: (2024)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
by: Song, Yuchen, et al.
Published: (2025)
by: Song, Yuchen, et al.
Published: (2025)
PEPR: Privileged Event-based Predictive Regularization for Domain Generalization
by: Magrini, Gabriele, et al.
Published: (2026)
by: Magrini, Gabriele, et al.
Published: (2026)
A benchmark dataset for deep learning-based airplane detection: HRPlanes
by: Bakirman, Tolga, et al.
Published: (2022)
by: Bakirman, Tolga, et al.
Published: (2022)
Backward-Compatible Aligned Representations via an Orthogonal Transformation Layer
by: Ricci, Simone, et al.
Published: (2024)
by: Ricci, Simone, et al.
Published: (2024)
Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model Replacements
by: Biondi, Niccolò, et al.
Published: (2024)
by: Biondi, Niccolò, et al.
Published: (2024)
Machine Unlearning for Document Classification
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
Similar Items
-
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
by: Vivoli, Emanuele, et al.
Published: (2024) -
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024) -
One missing piece in Vision and Language: A Survey on Comics Understanding
by: Vivoli, Emanuele, et al.
Published: (2024) -
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024) -
ComicsPAP: understanding comic strips by picking the correct panel
by: Vivoli, Emanuele, et al.
Published: (2025)