VideoDistill: Language-aware Vision Distillation for Video Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Bo, Yang, Chao, Qiao, Yu, Quan, Chengbin, Zhao, Youjian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
Distilling Vision-Language Models on Millions of Videos
von: Zhao, Yue, et al.
Veröffentlicht: (2024)
von: Zhao, Yue, et al.
Veröffentlicht: (2024)
ViLA: Efficient Video-Language Alignment for Video Question Answering
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
Dataset Distillation via Vision-Language Category Prototype
von: Zou, Yawen, et al.
Veröffentlicht: (2025)
von: Zou, Yawen, et al.
Veröffentlicht: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
Distill Video Datasets into Images
von: Zhao, Zhenghao, et al.
Veröffentlicht: (2025)
von: Zhao, Zhenghao, et al.
Veröffentlicht: (2025)
EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
von: Cai, Wenqi, et al.
Veröffentlicht: (2026)
von: Cai, Wenqi, et al.
Veröffentlicht: (2026)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
von: Bai, Xiangyu, et al.
Veröffentlicht: (2026)
von: Bai, Xiangyu, et al.
Veröffentlicht: (2026)
Agentic Keyframe Search for Video Question Answering
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
CPFD: Confidence-aware Privileged Feature Distillation for Short Video Classification
von: Shi, Jinghao, et al.
Veröffentlicht: (2024)
von: Shi, Jinghao, et al.
Veröffentlicht: (2024)
SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
von: Tang, Qi, et al.
Veröffentlicht: (2024)
von: Tang, Qi, et al.
Veröffentlicht: (2024)
Teeth-SEG: An Efficient Instance Segmentation Framework for Orthodontic Treatment based on Anthropic Prior Knowledge
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
Latent Video Dataset Distillation
von: Li, Ning, et al.
Veröffentlicht: (2025)
von: Li, Ning, et al.
Veröffentlicht: (2025)
Video Set Distillation: Information Diversification and Temporal Densification
von: Zhao, Yinjie, et al.
Veröffentlicht: (2024)
von: Zhao, Yinjie, et al.
Veröffentlicht: (2024)
PartDistill: 3D Shape Part Segmentation by Vision-Language Model Distillation
von: Umam, Ardian, et al.
Veröffentlicht: (2023)
von: Umam, Ardian, et al.
Veröffentlicht: (2023)
Training-Free Motion Customization for Distilled Video Generators with Adaptive Test-Time Distillation
von: Rong, Jintao, et al.
Veröffentlicht: (2025)
von: Rong, Jintao, et al.
Veröffentlicht: (2025)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
von: Nikandrou, Malvina, et al.
Veröffentlicht: (2024)
von: Nikandrou, Malvina, et al.
Veröffentlicht: (2024)
VidCtx: Context-aware Video Question Answering with Image Models
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
Vision-Language Dataset Distillation
von: Wu, Xindi, et al.
Veröffentlicht: (2023)
von: Wu, Xindi, et al.
Veröffentlicht: (2023)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
Streaming Autoregressive Video Generation via Diagonal Distillation
von: Liu, Jinxiu, et al.
Veröffentlicht: (2026)
von: Liu, Jinxiu, et al.
Veröffentlicht: (2026)
Narrative Aligned Long Form Video Question Answering
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
von: Zou, Bo, et al.
Veröffentlicht: (2024) -
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
von: Liang, Tianming, et al.
Veröffentlicht: (2024) -
Distilling Vision-Language Models on Millions of Videos
von: Zhao, Yue, et al.
Veröffentlicht: (2024) -
ViLA: Efficient Video-Language Alignment for Video Question Answering
von: Wang, Xijun, et al.
Veröffentlicht: (2023) -
Dataset Distillation via Vision-Language Category Prototype
von: Zou, Yawen, et al.
Veröffentlicht: (2025)