Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yingxuan, Hinami, Ryota, Aizawa, Kiyoharu, Matsui, Yusuke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
di: Li, Yingxuan, et al.
Pubblicazione: (2023)
di: Li, Yingxuan, et al.
Pubblicazione: (2023)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
di: Ikuta, Hikaru, et al.
Pubblicazione: (2024)
di: Ikuta, Hikaru, et al.
Pubblicazione: (2024)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
di: Imajuku, Yuki, et al.
Pubblicazione: (2024)
di: Imajuku, Yuki, et al.
Pubblicazione: (2024)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
di: Watanabe, Mitsuki, et al.
Pubblicazione: (2025)
di: Watanabe, Mitsuki, et al.
Pubblicazione: (2025)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
di: Liao, Liwei, et al.
Pubblicazione: (2025)
di: Liao, Liwei, et al.
Pubblicazione: (2025)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
di: Xu, Yifang, et al.
Pubblicazione: (2025)
di: Xu, Yifang, et al.
Pubblicazione: (2025)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
di: Zhao, Xinyu, et al.
Pubblicazione: (2026)
di: Zhao, Xinyu, et al.
Pubblicazione: (2026)
Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
di: Cai, Lingling, et al.
Pubblicazione: (2024)
di: Cai, Lingling, et al.
Pubblicazione: (2024)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
di: Lin, Haoqiang, et al.
Pubblicazione: (2025)
di: Lin, Haoqiang, et al.
Pubblicazione: (2025)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
di: Cheung, Tsun-Hin, et al.
Pubblicazione: (2024)
di: Cheung, Tsun-Hin, et al.
Pubblicazione: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
di: Duan, Yiqun, et al.
Pubblicazione: (2025)
di: Duan, Yiqun, et al.
Pubblicazione: (2025)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
di: Wen, Haokun, et al.
Pubblicazione: (2026)
di: Wen, Haokun, et al.
Pubblicazione: (2026)
MU-MAE: Multimodal Masked Autoencoders-Based One-Shot Learning
di: Liu, Rex, et al.
Pubblicazione: (2024)
di: Liu, Rex, et al.
Pubblicazione: (2024)
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
di: Li, Wenrui, et al.
Pubblicazione: (2024)
di: Li, Wenrui, et al.
Pubblicazione: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
di: Zhang, Beiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Beiyuan, et al.
Pubblicazione: (2024)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
di: Yang, Sicheng, et al.
Pubblicazione: (2025)
di: Yang, Sicheng, et al.
Pubblicazione: (2025)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
di: Zhu, Lingsi, et al.
Pubblicazione: (2026)
di: Zhu, Lingsi, et al.
Pubblicazione: (2026)
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
di: Tang, Hao, et al.
Pubblicazione: (2025)
di: Tang, Hao, et al.
Pubblicazione: (2025)
Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
di: Cai, Haonan, et al.
Pubblicazione: (2026)
di: Cai, Haonan, et al.
Pubblicazione: (2026)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
di: Yu, Hongyun, et al.
Pubblicazione: (2024)
di: Yu, Hongyun, et al.
Pubblicazione: (2024)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
di: Hoe, Jiun Tian, et al.
Pubblicazione: (2025)
di: Hoe, Jiun Tian, et al.
Pubblicazione: (2025)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
di: Zhang, Zhihao, et al.
Pubblicazione: (2023)
di: Zhang, Zhihao, et al.
Pubblicazione: (2023)
A Multimodal Transformer for Live Streaming Highlight Prediction
di: Deng, Jiaxin, et al.
Pubblicazione: (2024)
di: Deng, Jiaxin, et al.
Pubblicazione: (2024)
AGSP-DSA: An Adaptive Graph Signal Processing Framework for Robust Multimodal Fusion with Dynamic Semantic Alignment
di: Karthikeya, KV, et al.
Pubblicazione: (2026)
di: Karthikeya, KV, et al.
Pubblicazione: (2026)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
di: Wang, Zhenyu, et al.
Pubblicazione: (2026)
di: Wang, Zhenyu, et al.
Pubblicazione: (2026)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
di: Chen, Sen, et al.
Pubblicazione: (2022)
di: Chen, Sen, et al.
Pubblicazione: (2022)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026)
di: Song, Shezheng, et al.
Pubblicazione: (2026)
Gorgeous: Create Your Desired Character Facial Makeup from Any Ideas
di: Sii, Jia Wei, et al.
Pubblicazione: (2024)
di: Sii, Jia Wei, et al.
Pubblicazione: (2024)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
di: Wang, Xue, et al.
Pubblicazione: (2026)
di: Wang, Xue, et al.
Pubblicazione: (2026)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
di: Hsiao, Teng-Fang, et al.
Pubblicazione: (2024)
di: Hsiao, Teng-Fang, et al.
Pubblicazione: (2024)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
di: Liu, Han, et al.
Pubblicazione: (2025)
di: Liu, Han, et al.
Pubblicazione: (2025)
FakeBench: Probing Explainable Fake Image Detection via Large Multimodal Models
di: Li, Yixuan, et al.
Pubblicazione: (2024)
di: Li, Yixuan, et al.
Pubblicazione: (2024)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
di: Tan, Jiawei, et al.
Pubblicazione: (2024)
di: Tan, Jiawei, et al.
Pubblicazione: (2024)
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
di: Zhang, Yinghui, et al.
Pubblicazione: (2025)
di: Zhang, Yinghui, et al.
Pubblicazione: (2025)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
di: Zhu, Yixin, et al.
Pubblicazione: (2026)
di: Zhu, Yixin, et al.
Pubblicazione: (2026)
RED-DOT: Multimodal Fact-checking via Relevant Evidence Detection
di: Papadopoulos, Stefanos-Iordanis, et al.
Pubblicazione: (2023)
di: Papadopoulos, Stefanos-Iordanis, et al.
Pubblicazione: (2023)
Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
di: Sar, Ayan, et al.
Pubblicazione: (2025)
di: Sar, Ayan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
di: Li, Yingxuan, et al.
Pubblicazione: (2023) -
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
di: Ikuta, Hikaru, et al.
Pubblicazione: (2024) -
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
di: Imajuku, Yuki, et al.
Pubblicazione: (2024) -
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
di: Watanabe, Mitsuki, et al.
Pubblicazione: (2025) -
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
di: Liao, Liwei, et al.
Pubblicazione: (2025)