The Manga Whisperer: Automatically Generating Transcriptions for Comics
Fuente:
arXiv
Saved in:
| Main Authors: | Sachdeva, Ragav, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025)
by: Sachdeva, Ragav, et al.
Published: (2025)
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
by: Li, Yingxuan, et al.
Published: (2023)
by: Li, Yingxuan, et al.
Published: (2023)
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
by: Wang, Muyao, et al.
Published: (2026)
by: Wang, Muyao, et al.
Published: (2026)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025)
by: Korbar, Bruno, et al.
Published: (2025)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Manga Generation via Layout-controllable Diffusion
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Retrieval Augmented Comic Image Generation
by: Shui, Yunhao, et al.
Published: (2025)
by: Shui, Yunhao, et al.
Published: (2025)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
by: Ikuta, Hikaru, et al.
Published: (2024)
by: Ikuta, Hikaru, et al.
Published: (2024)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2026)
by: Baek, Jeonghun, et al.
Published: (2026)
Recognising BSL Fingerspelling in Continuous Signing Sequences
by: Chan, Alyssa, et al.
Published: (2026)
by: Chan, Alyssa, et al.
Published: (2026)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
Adapting MLLMs for Nuanced Video Retrieval
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024)
by: Amini-Naieni, Niki, et al.
Published: (2024)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
Moving Object Segmentation: All You Need Is SAM (and Flow)
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Recognizing Co-Speech Gestures in-the-Wild
by: Hegde, Sindhu B, et al.
Published: (2026)
by: Hegde, Sindhu B, et al.
Published: (2026)
Amodal Ground Truth and Completion in the Wild
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Number it: Temporal Grounding Videos like Flipping Manga
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
Inference-time Trajectory Optimization for Manga Image Editing
by: Furuta, Ryosuke
Published: (2026)
by: Furuta, Ryosuke
Published: (2026)
Made to Order: Discovering monotonic temporal changes via self-supervised video ordering
by: Yang, Charig, et al.
Published: (2024)
by: Yang, Charig, et al.
Published: (2024)
New keypoint-based approach for recognising British Sign Language (BSL) from sequences
by: Deb, Oishi, et al.
Published: (2024)
by: Deb, Oishi, et al.
Published: (2024)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Advanced Knowledge Extraction of Physical Design Drawings, Translation and conversion to CAD formats using Deep Learning
by: M, Jesher Joshua, et al.
Published: (2024)
by: M, Jesher Joshua, et al.
Published: (2024)
MangaNinja: Line Art Colorization with Precise Reference Following
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
CTC Transcription Alignment of the Bullinger Letters: Automatic Improvement of Annotation Quality
by: Peer, Marco, et al.
Published: (2025)
by: Peer, Marco, et al.
Published: (2025)
Region-Wise Correspondence Prediction between Manga Line Art Images
by: Li, Yingxuan, et al.
Published: (2025)
by: Li, Yingxuan, et al.
Published: (2025)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
by: Chalk, Jacob, et al.
Published: (2024)
by: Chalk, Jacob, et al.
Published: (2024)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
Understanding Co-speech Gestures in-the-wild
by: Hegde, Sindhu B, et al.
Published: (2025)
by: Hegde, Sindhu B, et al.
Published: (2025)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
by: Jang, Youngjoon, et al.
Published: (2025)
by: Jang, Youngjoon, et al.
Published: (2025)
Similar Items
-
Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names
by: Sachdeva, Ragav, et al.
Published: (2024) -
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025) -
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
by: Li, Yingxuan, et al.
Published: (2023) -
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025) -
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
by: Wang, Muyao, et al.
Published: (2026)