Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yingxuan, Aizawa, Kiyoharu, Matsui, Yusuke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
di: Li, Yingxuan, et al.
Pubblicazione: (2024)
di: Li, Yingxuan, et al.
Pubblicazione: (2024)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
di: Baek, Jeonghun, et al.
Pubblicazione: (2026)
di: Baek, Jeonghun, et al.
Pubblicazione: (2026)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
di: Ikuta, Hikaru, et al.
Pubblicazione: (2024)
di: Ikuta, Hikaru, et al.
Pubblicazione: (2024)
Region-Wise Correspondence Prediction between Manga Line Art Images
di: Li, Yingxuan, et al.
Pubblicazione: (2025)
di: Li, Yingxuan, et al.
Pubblicazione: (2025)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
di: Baek, Jeonghun, et al.
Pubblicazione: (2025)
di: Baek, Jeonghun, et al.
Pubblicazione: (2025)
Noisy Label Refinement with Semantically Reliable Synthetic Images
di: Li, Yingxuan, et al.
Pubblicazione: (2025)
di: Li, Yingxuan, et al.
Pubblicazione: (2025)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
di: Baek, Jeonghun, et al.
Pubblicazione: (2025)
di: Baek, Jeonghun, et al.
Pubblicazione: (2025)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
di: Sachdeva, Ragav, et al.
Pubblicazione: (2024)
di: Sachdeva, Ragav, et al.
Pubblicazione: (2024)
Entity-NeRF: Detecting and Removing Moving Entities in Urban Scenes
di: Otonari, Takashi, et al.
Pubblicazione: (2024)
di: Otonari, Takashi, et al.
Pubblicazione: (2024)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
di: Imajuku, Yuki, et al.
Pubblicazione: (2024)
di: Imajuku, Yuki, et al.
Pubblicazione: (2024)
GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2023)
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2023)
Guided Image Synthesis via Initial Image Editing in Diffusion Model
di: Mao, Jiafeng, et al.
Pubblicazione: (2023)
di: Mao, Jiafeng, et al.
Pubblicazione: (2023)
The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
di: Mao, Jiafeng, et al.
Pubblicazione: (2023)
di: Mao, Jiafeng, et al.
Pubblicazione: (2023)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
di: Watanabe, Mitsuki, et al.
Pubblicazione: (2025)
di: Watanabe, Mitsuki, et al.
Pubblicazione: (2025)
A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
di: Noda, Shiho, et al.
Pubblicazione: (2025)
di: Noda, Shiho, et al.
Pubblicazione: (2025)
Training-Free Sketch-Guided Diffusion with Latent Optimization
di: Ding, Sandra Zhang, et al.
Pubblicazione: (2024)
di: Ding, Sandra Zhang, et al.
Pubblicazione: (2024)
CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains
di: Takeda, Tomohisa, et al.
Pubblicazione: (2026)
di: Takeda, Tomohisa, et al.
Pubblicazione: (2026)
PerFace: Metric Learning in Perceptual Facial Similarity for Enhanced Face Anonymization
di: Kumagai, Haruka, et al.
Pubblicazione: (2025)
di: Kumagai, Haruka, et al.
Pubblicazione: (2025)
SVGEditBench: A Benchmark Dataset for Quantitative Assessment of LLM's SVG Editing Capabilities
di: Nishina, Kunato, et al.
Pubblicazione: (2024)
di: Nishina, Kunato, et al.
Pubblicazione: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
di: Vivoli, Emanuele, et al.
Pubblicazione: (2024)
di: Vivoli, Emanuele, et al.
Pubblicazione: (2024)
Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
di: Horita, Daichi, et al.
Pubblicazione: (2023)
di: Horita, Daichi, et al.
Pubblicazione: (2023)
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2025)
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2025)
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset
di: Lee, Young-Jun, et al.
Pubblicazione: (2022)
di: Lee, Young-Jun, et al.
Pubblicazione: (2022)
RouteExtract: A Modular Pipeline for Extracting Routes from Paper Maps
di: Kremser, Bjoern, et al.
Pubblicazione: (2025)
di: Kremser, Bjoern, et al.
Pubblicazione: (2025)
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
di: Wang, Muyao, et al.
Pubblicazione: (2026)
di: Wang, Muyao, et al.
Pubblicazione: (2026)
Adversarial Doodles: Interpretable and Human-drawable Attacks Provide Describable Insights
di: Nara, Ryoya, et al.
Pubblicazione: (2023)
di: Nara, Ryoya, et al.
Pubblicazione: (2023)
High-Frequency Anti-DreamBooth: Robust Defense against Personalized Image Synthesis
di: Onikubo, Takuto, et al.
Pubblicazione: (2024)
di: Onikubo, Takuto, et al.
Pubblicazione: (2024)
Manga Generation via Layout-controllable Diffusion
di: Chen, Siyu, et al.
Pubblicazione: (2024)
di: Chen, Siyu, et al.
Pubblicazione: (2024)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2024)
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2024)
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
di: Katsube, Toshiki, et al.
Pubblicazione: (2025)
di: Katsube, Toshiki, et al.
Pubblicazione: (2025)
LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff Table
di: Matsui, Yusuke
Pubblicazione: (2025)
di: Matsui, Yusuke
Pubblicazione: (2025)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2025)
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2025)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
di: Huang, Minbin, et al.
Pubblicazione: (2024)
di: Huang, Minbin, et al.
Pubblicazione: (2024)
ZoDi: Zero-Shot Domain Adaptation with Diffusion-Based Image Transfer
di: Azuma, Hiroki, et al.
Pubblicazione: (2024)
di: Azuma, Hiroki, et al.
Pubblicazione: (2024)
Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
Retrieval Augmented Comic Image Generation
di: Shui, Yunhao, et al.
Pubblicazione: (2025)
di: Shui, Yunhao, et al.
Pubblicazione: (2025)
Alignment-Free RGB-T Salient Object Detection: A Large-scale Dataset and Progressive Correlation Network
di: Wang, Kunpeng, et al.
Pubblicazione: (2024)
di: Wang, Kunpeng, et al.
Pubblicazione: (2024)
Unlocking Comics: The AI4VA Dataset for Visual Understanding
di: Grönquist, Peter, et al.
Pubblicazione: (2024)
di: Grönquist, Peter, et al.
Pubblicazione: (2024)
WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection
di: Hong, Yan, et al.
Pubblicazione: (2024)
di: Hong, Yan, et al.
Pubblicazione: (2024)
Inference-time Trajectory Optimization for Manga Image Editing
di: Furuta, Ryosuke
Pubblicazione: (2026)
di: Furuta, Ryosuke
Pubblicazione: (2026)
Documenti analoghi
-
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
di: Li, Yingxuan, et al.
Pubblicazione: (2024) -
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
di: Baek, Jeonghun, et al.
Pubblicazione: (2026) -
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
di: Ikuta, Hikaru, et al.
Pubblicazione: (2024) -
Region-Wise Correspondence Prediction between Manga Line Art Images
di: Li, Yingxuan, et al.
Pubblicazione: (2025) -
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
di: Baek, Jeonghun, et al.
Pubblicazione: (2025)