CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ortega, Marc Serra, Vivoli, Emanuele, Llabrés, Artemis, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
ComicsPAP: understanding comic strips by picking the correct panel
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
Multimodal Transformer for Comics Text-Cloze
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
One missing piece in Vision and Language: A Survey on Comics Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)
Image-text matching for large-scale book collections
von: Llabrés, Artemis, et al.
Veröffentlicht: (2024)
von: Llabrés, Artemis, et al.
Veröffentlicht: (2024)
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
von: Jin, Li, et al.
Veröffentlicht: (2026)
von: Jin, Li, et al.
Veröffentlicht: (2026)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
von: Kang, Lei, et al.
Veröffentlicht: (2024)
von: Kang, Lei, et al.
Veröffentlicht: (2024)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
von: Gomez, Lluis, et al.
Veröffentlicht: (2014)
von: Gomez, Lluis, et al.
Veröffentlicht: (2014)
Enhancing Document VQA Models via Retrieval-Augmented Generation
von: López, Eric, et al.
Veröffentlicht: (2025)
von: López, Eric, et al.
Veröffentlicht: (2025)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
von: Li, Zongmin, et al.
Veröffentlicht: (2026)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
von: Mishra, Pritam, et al.
Veröffentlicht: (2025)
von: Mishra, Pritam, et al.
Veröffentlicht: (2025)
Federated Document Visual Question Answering: A Pilot Study
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
von: Mishra, Pritam, et al.
Veröffentlicht: (2026)
von: Mishra, Pritam, et al.
Veröffentlicht: (2026)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
von: Chattopadhyay, Soumitri, et al.
Veröffentlicht: (2024)
von: Chattopadhyay, Soumitri, et al.
Veröffentlicht: (2024)
Reading in the Dark: Low-light Scene Text Recognition
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
von: Fu, Xuanshuo, et al.
Veröffentlicht: (2026)
Reading Between the Lanes: Text VideoQA on the Road
von: Tom, George, et al.
Veröffentlicht: (2023)
von: Tom, George, et al.
Veröffentlicht: (2023)
MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow
von: Li, Zhe, et al.
Veröffentlicht: (2024)
von: Li, Zhe, et al.
Veröffentlicht: (2024)
HoloMine: A Synthetic Dataset for Buried Landmines Recognition using Microwave Holographic Imaging
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
von: Kang, Lei, et al.
Veröffentlicht: (2025)
von: Kang, Lei, et al.
Veröffentlicht: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation
von: Dey, Arka Ujjal, et al.
Veröffentlicht: (2024)
von: Dey, Arka Ujjal, et al.
Veröffentlicht: (2024)
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts
von: Wu, Jialin, et al.
Veröffentlicht: (2023)
von: Wu, Jialin, et al.
Veröffentlicht: (2023)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
Machine Unlearning for Document Classification
von: Kang, Lei, et al.
Veröffentlicht: (2024)
von: Kang, Lei, et al.
Veröffentlicht: (2024)
BookNet: Book Image Rectification via Cross-Page Attention Network
von: Liu, Shaokai, et al.
Veröffentlicht: (2026)
von: Liu, Shaokai, et al.
Veröffentlicht: (2026)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
A Multimodal Transformer for Live Streaming Highlight Prediction
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
Retrieval Augmented Comic Image Generation
von: Shui, Yunhao, et al.
Veröffentlicht: (2025)
von: Shui, Yunhao, et al.
Veröffentlicht: (2025)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
von: Sachdeva, Ragav, et al.
Veröffentlicht: (2024)
GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models
von: Kang, Lei, et al.
Veröffentlicht: (2024)
von: Kang, Lei, et al.
Veröffentlicht: (2024)
Labeling Comic Mischief Content in Online Videos with a Multimodal Hierarchical-Cross-Attention Model
von: Baharlouei, Elaheh, et al.
Veröffentlicht: (2024)
von: Baharlouei, Elaheh, et al.
Veröffentlicht: (2024)
TextBite: A Historical Czech Document Dataset for Logical Page Segmentation
von: Kostelník, Martin, et al.
Veröffentlicht: (2025)
von: Kostelník, Martin, et al.
Veröffentlicht: (2025)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
Dual-Stream Alignment for Action Segmentation
von: Gammulle, Harshala, et al.
Veröffentlicht: (2025)
von: Gammulle, Harshala, et al.
Veröffentlicht: (2025)
USCNet: Transformer-Based Multimodal Fusion with Segmentation Guidance for Urolithiasis Classification
von: Wang, Changmiao, et al.
Veröffentlicht: (2026)
von: Wang, Changmiao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024) -
ComicsPAP: understanding comic strips by picking the correct panel
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2025) -
Multimodal Transformer for Comics Text-Cloze
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024) -
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024) -
One missing piece in Vision and Language: A Survey on Comics Understanding
von: Vivoli, Emanuele, et al.
Veröffentlicht: (2024)