ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phukan, Arpan, Gupta, Manish, Ekbal, Asif |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
von: Phukan, Arpan, et al.
Veröffentlicht: (2025)
von: Phukan, Arpan, et al.
Veröffentlicht: (2025)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
von: Pandey, Anupam, et al.
Veröffentlicht: (2025)
von: Pandey, Anupam, et al.
Veröffentlicht: (2025)
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages
von: Gain, Baban, et al.
Veröffentlicht: (2023)
von: Gain, Baban, et al.
Veröffentlicht: (2023)
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
SparrowVQE: Visual Question Explanation for Course Content Understanding
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Vision-centric Token Compression in Large Language Model
von: Xing, Ling, et al.
Veröffentlicht: (2025)
von: Xing, Ling, et al.
Veröffentlicht: (2025)
Actions and Objects Pathways for Domain Adaptation in Video Question Answering
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
HARE: an entity and relation centric evaluation framework for histopathology reports
von: Kim, Yunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2025)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
von: Stamatakis, Markos, et al.
Veröffentlicht: (2025)
von: Stamatakis, Markos, et al.
Veröffentlicht: (2025)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
von: Cai, Chen, et al.
Veröffentlicht: (2024)
von: Cai, Chen, et al.
Veröffentlicht: (2024)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
WikiVideo: Article Generation from Multiple Videos
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
Generating Fine Details of Entity Interactions
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
Video Summarization: Towards Entity-Aware Captions
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
von: Ayyubi, Hammad A., et al.
Veröffentlicht: (2023)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language Model
von: Kim, Taehee, et al.
Veröffentlicht: (2024)
von: Kim, Taehee, et al.
Veröffentlicht: (2024)
Grounding Language Models for Visual Entity Recognition
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning
von: Zhang, Peifeng, et al.
Veröffentlicht: (2026)
von: Zhang, Peifeng, et al.
Veröffentlicht: (2026)
What is Beneath Misogyny: Misogynous Memes Classification and Explanation
von: Kanwar, Kushal, et al.
Veröffentlicht: (2025)
von: Kanwar, Kushal, et al.
Veröffentlicht: (2025)
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
von: Wu, Shu, et al.
Veröffentlicht: (2026)
von: Wu, Shu, et al.
Veröffentlicht: (2026)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
von: Shen, Kai, et al.
Veröffentlicht: (2024)
von: Shen, Kai, et al.
Veröffentlicht: (2024)
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
von: Long, Fuchen, et al.
Veröffentlicht: (2024)
von: Long, Fuchen, et al.
Veröffentlicht: (2024)
Generalizable Entity Grounding via Assistance of Large Language Model
von: Qi, Lu, et al.
Veröffentlicht: (2024)
von: Qi, Lu, et al.
Veröffentlicht: (2024)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
Text-centric Alignment for Multi-Modality Learning
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports
von: Serra, Francesco Dalla, et al.
Veröffentlicht: (2025)
von: Serra, Francesco Dalla, et al.
Veröffentlicht: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
von: Xu, Jilan, et al.
Veröffentlicht: (2025)
von: Xu, Jilan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
von: Phukan, Arpan, et al.
Veröffentlicht: (2025) -
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
von: Pandey, Anupam, et al.
Veröffentlicht: (2025) -
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
von: Mi, Li, et al.
Veröffentlicht: (2024) -
Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages
von: Gain, Baban, et al.
Veröffentlicht: (2023) -
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)