A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
Fuente:
arXiv
Guardado en:
| Autores principales: | Imrattanatrai, Wiradee, Asada, Masaki, Hasegawa, Kimihiro, Cheng, Zhi-Qi, Fukuda, Ken, Mitamura, Teruko |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
por: Hasegawa, Kimihiro, et al.
Publicado: (2025)
por: Hasegawa, Kimihiro, et al.
Publicado: (2025)
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
por: Hasegawa, Kimihiro, et al.
Publicado: (2025)
por: Hasegawa, Kimihiro, et al.
Publicado: (2025)
ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
por: Hasegawa, Kimihiro, et al.
Publicado: (2024)
por: Hasegawa, Kimihiro, et al.
Publicado: (2024)
Formulation Comparison for Timeline Construction using LLMs
por: Hasegawa, Kimihiro, et al.
Publicado: (2024)
por: Hasegawa, Kimihiro, et al.
Publicado: (2024)
Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches
por: Pratapa, Adithya, et al.
Publicado: (2025)
por: Pratapa, Adithya, et al.
Publicado: (2025)
Estimating Optimal Context Length for Hybrid Retrieval-augmented Multi-document Summarization
por: Pratapa, Adithya, et al.
Publicado: (2025)
por: Pratapa, Adithya, et al.
Publicado: (2025)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
por: Ikoma, Hayato, et al.
Publicado: (2025)
por: Ikoma, Hayato, et al.
Publicado: (2025)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
por: Reichman, Benjamin, et al.
Publicado: (2025)
por: Reichman, Benjamin, et al.
Publicado: (2025)
Video sentence grounding with temporally global textual knowledge
por: Chen, Cai, et al.
Publicado: (2024)
por: Chen, Cai, et al.
Publicado: (2024)
VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
por: Glória-Silva, Diogo, et al.
Publicado: (2026)
por: Glória-Silva, Diogo, et al.
Publicado: (2026)
VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
por: Cheng, Ming, et al.
Publicado: (2025)
por: Cheng, Ming, et al.
Publicado: (2025)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
por: Yoon, Eunseop, et al.
Publicado: (2025)
por: Yoon, Eunseop, et al.
Publicado: (2025)
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset
por: Lee, Young-Jun, et al.
Publicado: (2022)
por: Lee, Young-Jun, et al.
Publicado: (2022)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
por: Kim, Youngmin, et al.
Publicado: (2025)
por: Kim, Youngmin, et al.
Publicado: (2025)
LLM-grounded Video Diffusion Models
por: Lian, Long, et al.
Publicado: (2023)
por: Lian, Long, et al.
Publicado: (2023)
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
por: Inadumi, Shun, et al.
Publicado: (2024)
por: Inadumi, Shun, et al.
Publicado: (2024)
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
por: Peng, Bo, et al.
Publicado: (2026)
por: Peng, Bo, et al.
Publicado: (2026)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
por: Li, Heng, et al.
Publicado: (2024)
por: Li, Heng, et al.
Publicado: (2024)
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
por: Wang, Yan, et al.
Publicado: (2024)
por: Wang, Yan, et al.
Publicado: (2024)
Towards Event-oriented Long Video Understanding
por: Du, Yifan, et al.
Publicado: (2024)
por: Du, Yifan, et al.
Publicado: (2024)
VHAKG: A Multi-modal Knowledge Graph Based on Synchronized Multi-view Videos of Daily Activities
por: Egami, Shusaku, et al.
Publicado: (2024)
por: Egami, Shusaku, et al.
Publicado: (2024)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
por: Liang, Hao, et al.
Publicado: (2024)
por: Liang, Hao, et al.
Publicado: (2024)
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
por: Biyyala, Varun, et al.
Publicado: (2025)
por: Biyyala, Varun, et al.
Publicado: (2025)
RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events
por: Chen, Zhenyuan, et al.
Publicado: (2025)
por: Chen, Zhenyuan, et al.
Publicado: (2025)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
por: Liang, Jiafeng, et al.
Publicado: (2024)
por: Liang, Jiafeng, et al.
Publicado: (2024)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
por: Tong, Tony Cheng, et al.
Publicado: (2024)
por: Tong, Tony Cheng, et al.
Publicado: (2024)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
por: Tang, Yolo Yunlong, et al.
Publicado: (2023)
por: Tang, Yolo Yunlong, et al.
Publicado: (2023)
Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
por: Chew, Oscar, et al.
Publicado: (2026)
por: Chew, Oscar, et al.
Publicado: (2026)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
por: Ohkawa, Takehiko, et al.
Publicado: (2023)
por: Ohkawa, Takehiko, et al.
Publicado: (2023)
Hummus: A Dataset of Humorous Multimodal Metaphor Use
por: Tong, Xiaoyu, et al.
Publicado: (2025)
por: Tong, Xiaoyu, et al.
Publicado: (2025)
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
por: Kriz, Reno, et al.
Publicado: (2024)
por: Kriz, Reno, et al.
Publicado: (2024)
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
Towards Patronizing and Condescending Language in Chinese Videos: A Multimodal Dataset and Detector
por: Wang, Hongbo, et al.
Publicado: (2024)
por: Wang, Hongbo, et al.
Publicado: (2024)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
por: Chia, Yew Ken, et al.
Publicado: (2024)
por: Chia, Yew Ken, et al.
Publicado: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
por: Zhang, Xueqiao, et al.
Publicado: (2025)
por: Zhang, Xueqiao, et al.
Publicado: (2025)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
por: Mu, Kuan-Chen, et al.
Publicado: (2024)
por: Mu, Kuan-Chen, et al.
Publicado: (2024)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
Fostering Video Reasoning via Next-Event Prediction
por: Wang, Haonan, et al.
Publicado: (2025)
por: Wang, Haonan, et al.
Publicado: (2025)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
por: Han, Songhao, et al.
Publicado: (2024)
por: Han, Songhao, et al.
Publicado: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
por: Zang, Yuan, et al.
Publicado: (2025)
por: Zang, Yuan, et al.
Publicado: (2025)
Ejemplares similares
-
TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
por: Hasegawa, Kimihiro, et al.
Publicado: (2025) -
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
por: Hasegawa, Kimihiro, et al.
Publicado: (2025) -
ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
por: Hasegawa, Kimihiro, et al.
Publicado: (2024) -
Formulation Comparison for Timeline Construction using LLMs
por: Hasegawa, Kimihiro, et al.
Publicado: (2024) -
Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches
por: Pratapa, Adithya, et al.
Publicado: (2025)