ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy
Fuente:
arXiv
Guardado en:
| Autores principales: | Kawamura, Kazuki, Nakai, Kengo, Rekimoto, Jun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Contexts
por: Kawamura, Kazuki, et al.
Publicado: (2024)
por: Kawamura, Kazuki, et al.
Publicado: (2024)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
por: Imajuku, Yuki, et al.
Publicado: (2024)
por: Imajuku, Yuki, et al.
Publicado: (2024)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
por: Yu, Lei, et al.
Publicado: (2024)
por: Yu, Lei, et al.
Publicado: (2024)
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
por: Chen, Lizhi, et al.
Publicado: (2025)
por: Chen, Lizhi, et al.
Publicado: (2025)
"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
por: Skoularikis, Anastasios, et al.
Publicado: (2025)
por: Skoularikis, Anastasios, et al.
Publicado: (2025)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
por: Chu, Meng, et al.
Publicado: (2025)
por: Chu, Meng, et al.
Publicado: (2025)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
por: Xu, Shuolin, et al.
Publicado: (2025)
por: Xu, Shuolin, et al.
Publicado: (2025)
Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
por: Hao, Jing, et al.
Publicado: (2025)
por: Hao, Jing, et al.
Publicado: (2025)
PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
por: Goncalves, Lucas, et al.
Publicado: (2024)
por: Goncalves, Lucas, et al.
Publicado: (2024)
Efficiently Collecting Training Dataset for 2D Object Detection by Online Visual Feedback
por: Kiyokawa, Takuya, et al.
Publicado: (2023)
por: Kiyokawa, Takuya, et al.
Publicado: (2023)
3D2M Dataset: A 3-Dimension diverse Mesh Dataset
por: Dasgupta, Sankarshan
Publicado: (2024)
por: Dasgupta, Sankarshan
Publicado: (2024)
SakugaFlow: A Stagewise Illustration Framework Emulating the Human Drawing Process and Providing Interactive Tutoring for Novice Drawing Skills
por: Kawamura, Kazuki, et al.
Publicado: (2025)
por: Kawamura, Kazuki, et al.
Publicado: (2025)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
por: Zhu, Lingsi, et al.
Publicado: (2026)
por: Zhu, Lingsi, et al.
Publicado: (2026)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
por: Zhang, Xueqiao, et al.
Publicado: (2025)
por: Zhang, Xueqiao, et al.
Publicado: (2025)
CrypticBio: A Large Multimodal Dataset for Visually Confusing Biodiversity
por: Manolache, Georgiana, et al.
Publicado: (2025)
por: Manolache, Georgiana, et al.
Publicado: (2025)
A Multimodal Transformer for Live Streaming Highlight Prediction
por: Deng, Jiaxin, et al.
Publicado: (2024)
por: Deng, Jiaxin, et al.
Publicado: (2024)
GLENDA: Gynecologic Laparoscopy Endometriosis Dataset
por: Leibetseder, Andreas, et al.
Publicado: (2025)
por: Leibetseder, Andreas, et al.
Publicado: (2025)
Detached and Interactive Multimodal Learning
por: Fan, Yunfeng, et al.
Publicado: (2024)
por: Fan, Yunfeng, et al.
Publicado: (2024)
LMVD: A Large-Scale Multimodal Vlog Dataset for Depression Detection in the Wild
por: He, Lang, et al.
Publicado: (2024)
por: He, Lang, et al.
Publicado: (2024)
FMNV: A Dataset of Media-Published News Videos for Fake News Detection
por: Wang, Yihao, et al.
Publicado: (2025)
por: Wang, Yihao, et al.
Publicado: (2025)
SEED: A Benchmark Dataset for Sequential Facial Attribute Editing with Diffusion Models
por: Zhu, Yule, et al.
Publicado: (2025)
por: Zhu, Yule, et al.
Publicado: (2025)
GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality
por: Zhang, Zhiwei, et al.
Publicado: (2026)
por: Zhang, Zhiwei, et al.
Publicado: (2026)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
por: Ikuta, Hikaru, et al.
Publicado: (2024)
por: Ikuta, Hikaru, et al.
Publicado: (2024)
Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline
por: Wu, Xuecheng, et al.
Publicado: (2023)
por: Wu, Xuecheng, et al.
Publicado: (2023)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
por: Hao, Jing, et al.
Publicado: (2025)
por: Hao, Jing, et al.
Publicado: (2025)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
por: Chen, Shuang, et al.
Publicado: (2026)
por: Chen, Shuang, et al.
Publicado: (2026)
A Novel Multimodal System to Predict Agitation in People with Dementia Within Clinical Settings: A Proof of Concept
por: Badawi, Abeer, et al.
Publicado: (2024)
por: Badawi, Abeer, et al.
Publicado: (2024)
Learning Video Context as Interleaved Multimodal Sequences
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
por: Lin, Kevin Qinghong, et al.
Publicado: (2024)
A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods
por: Kontostathis, Ioannis, et al.
Publicado: (2024)
por: Kontostathis, Ioannis, et al.
Publicado: (2024)
When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
por: Gonzálbez-Biosca, Daniel, et al.
Publicado: (2025)
por: Gonzálbez-Biosca, Daniel, et al.
Publicado: (2025)
Scaling Audio-Visual Quality Assessment Dataset via Crowdsourcing
por: Yang, Renyu, et al.
Publicado: (2026)
por: Yang, Renyu, et al.
Publicado: (2026)
Grounded Chain-of-Thought for Multimodal Large Language Models
por: Wu, Qiong, et al.
Publicado: (2025)
por: Wu, Qiong, et al.
Publicado: (2025)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
por: Pramanick, Shraman, et al.
Publicado: (2025)
por: Pramanick, Shraman, et al.
Publicado: (2025)
Multimodal Engagement Analysis from Facial Videos in the Classroom
por: Sümer, Ömer, et al.
Publicado: (2021)
por: Sümer, Ömer, et al.
Publicado: (2021)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
por: Jiang, Jingjing, et al.
Publicado: (2025)
por: Jiang, Jingjing, et al.
Publicado: (2025)
MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
por: Zhao, Haochen, et al.
Publicado: (2025)
por: Zhao, Haochen, et al.
Publicado: (2025)
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
por: Ancarani, Elisa, et al.
Publicado: (2025)
por: Ancarani, Elisa, et al.
Publicado: (2025)
Latent Reconstruction from Generated Data for Multimodal Misinformation Detection
por: Papadopoulos, Stefanos-Iordanis, et al.
Publicado: (2025)
por: Papadopoulos, Stefanos-Iordanis, et al.
Publicado: (2025)
Can Multimodal Large Language Models Understand Spatial Relations?
por: Liu, Jingping, et al.
Publicado: (2025)
por: Liu, Jingping, et al.
Publicado: (2025)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
por: Zhou, Hengyang, et al.
Publicado: (2025)
por: Zhou, Hengyang, et al.
Publicado: (2025)
Ejemplares similares
-
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Contexts
por: Kawamura, Kazuki, et al.
Publicado: (2024) -
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
por: Imajuku, Yuki, et al.
Publicado: (2024) -
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
por: Yu, Lei, et al.
Publicado: (2024) -
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
por: Chen, Lizhi, et al.
Publicado: (2025) -
"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
por: Skoularikis, Anastasios, et al.
Publicado: (2025)