Question-Answering Dense Video Events
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Hangyu, Xiao, Junbin, Yao, Angela |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
VidCtx: Context-aware Video Question Answering with Image Models
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
von: Han, Wei, et al.
Veröffentlicht: (2023)
von: Han, Wei, et al.
Veröffentlicht: (2023)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Audio-visual Event Localization on Portrait Mode Short Videos
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
PolySmart @ TRECVid 2024 Medical Video Question Answering
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
von: Qing, Yuan, et al.
Veröffentlicht: (2026)
von: Qing, Yuan, et al.
Veröffentlicht: (2026)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2022)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2022)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
von: Qin, You, et al.
Veröffentlicht: (2024)
von: Qin, You, et al.
Veröffentlicht: (2024)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
von: Xiao, Xinyu, et al.
Veröffentlicht: (2026)
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
von: Qu, Qiang, et al.
Veröffentlicht: (2024)
Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
von: Yu, Zongyou, et al.
Veröffentlicht: (2024)
von: Yu, Zongyou, et al.
Veröffentlicht: (2024)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
Video Seal: Open and Efficient Video Watermarking
von: Fernandez, Pierre, et al.
Veröffentlicht: (2024)
von: Fernandez, Pierre, et al.
Veröffentlicht: (2024)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024)
von: Gu, Jing, et al.
Veröffentlicht: (2024)
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
von: Dahal, Ashim, et al.
Veröffentlicht: (2025)
von: Dahal, Ashim, et al.
Veröffentlicht: (2025)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023) -
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026) -
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024) -
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025) -
VidCtx: Context-aware Video Question Answering with Image Models
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)