ComiCap: A VLMs pipeline for dense captioning of Comic Panels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vivoli, Emanuele, Biondi, Niccolò, Bertini, Marco, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
One missing piece in Vision and Language: A Survey on Comics Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
ComicsPAP: understanding comic strips by picking the correct panel
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
von: Ortega, Marc Serra, et al.
Veröffentlicht: (2025)
von: Ortega, Marc Serra, et al.
Veröffentlicht: (2025)
Multimodal Transformer for Comics Text-Cloze
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
HoloMine: A Synthetic Dataset for Buried Landmines Recognition using Microwave Holographic Imaging
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
von: Gomez, Lluis, et al.
Veröffentlicht: (2014)
von: Gomez, Lluis, et al.
Veröffentlicht: (2014)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
von: Mishra, Pritam, et al.
Veröffentlicht: (2025)
von: Mishra, Pritam, et al.
Veröffentlicht: (2025)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
Federated Document Visual Question Answering: A Pilot Study
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
von: Mishra, Pritam, et al.
Veröffentlicht: (2026)
von: Mishra, Pritam, et al.
Veröffentlicht: (2026)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
von: Kang, Lei, et al.
Veröffentlicht: (2024)
von: Kang, Lei, et al.
Veröffentlicht: (2024)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
von: Chattopadhyay, Soumitri, et al.
Veröffentlicht: (2024)
von: Chattopadhyay, Soumitri, et al.
Veröffentlicht: (2024)
From Panels to Prose: Generating Literary Narratives from Comics
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2025)
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2025)
Reading in the Dark: Low-light Scene Text Recognition
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
Reading Between the Lanes: Text VideoQA on the Road
von: Tom, George, et al.
Veröffentlicht: (2023)
von: Tom, George, et al.
Veröffentlicht: (2023)
Image-text matching for large-scale book collections
von: Llabrés, Artemis, et al.
Veröffentlicht: (2024)
von: Llabrés, Artemis, et al.
Veröffentlicht: (2024)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
von: Kang, Lei, et al.
Veröffentlicht: (2025)
von: Kang, Lei, et al.
Veröffentlicht: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
Image captioning in different languages
von: van Miltenburg, Emiel
Veröffentlicht: (2024)
von: van Miltenburg, Emiel
Veröffentlicht: (2024)
PEPR: Privileged Event-based Predictive Regularization for Domain Generalization
von: Magrini, Gabriele, et al.
Veröffentlicht: (2026)
von: Magrini, Gabriele, et al.
Veröffentlicht: (2026)
Backward-Compatible Aligned Representations via an Orthogonal Transformation Layer
von: Ricci, Simone, et al.
Veröffentlicht: (2024)
von: Ricci, Simone, et al.
Veröffentlicht: (2024)
Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model Replacements
von: Biondi, Niccolò, et al.
Veröffentlicht: (2024)
von: Biondi, Niccolò, et al.
Veröffentlicht: (2024)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
Machine Unlearning for Document Classification
von: Kang, Lei, et al.
Veröffentlicht: (2024)
von: Kang, Lei, et al.
Veröffentlicht: (2024)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
Leveraging image captions for selective whole slide image annotation
von: Qiu, Jingna, et al.
Veröffentlicht: (2024)
von: Qiu, Jingna, et al.
Veröffentlicht: (2024)
Fine-grained length controllable video captioning with ordinal embeddings
von: Nitta, Tomoya, et al.
Veröffentlicht: (2024)
von: Nitta, Tomoya, et al.
Veröffentlicht: (2024)
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
von: Krestenitis, Marios, et al.
Veröffentlicht: (2026)
von: Krestenitis, Marios, et al.
Veröffentlicht: (2026)
FRED: The Florence RGB-Event Drone Dataset
von: Magrini, Gabriele, et al.
Veröffentlicht: (2025)
von: Magrini, Gabriele, et al.
Veröffentlicht: (2025)
Multi-Modal interpretable automatic video captioning
von: Hanna-Asaad, Antoine, et al.
Veröffentlicht: (2024)
von: Hanna-Asaad, Antoine, et al.
Veröffentlicht: (2024)
Mitigating Negative Flips via Margin Preserving Training
von: Ricci, Simone, et al.
Veröffentlicht: (2025)
von: Ricci, Simone, et al.
Veröffentlicht: (2025)
Object-oriented backdoor attack against image captioning
von: Li, Meiling, et al.
Veröffentlicht: (2024)
von: Li, Meiling, et al.
Veröffentlicht: (2024)
Retrieval Augmented Comic Image Generation
von: Shui, Yunhao, et al.
Veröffentlicht: (2025)
von: Shui, Yunhao, et al.
Veröffentlicht: (2025)
Image captioning for Brazilian Portuguese using GRIT model
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
Improving face generation quality and prompt following with synthetic captions
von: Tarasiou, Michail, et al.
Veröffentlicht: (2024)
von: Tarasiou, Michail, et al.
Veröffentlicht: (2024)
Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
von: Agnolucci, Lorenzo, et al.
Veröffentlicht: (2024)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024) -
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024) -
One missing piece in Vision and Language: A Survey on Comics Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024) -
ComicsPAP: understanding comic strips by picking the correct panel
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025) -
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
von: Ortega, Marc Serra, et al.
Veröffentlicht: (2025)