DrVideo: Document Retrieval Based Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Ziyu, Gou, Chenhui, Shi, Hengcan, Sun, Bin, Li, Shutao, Rezatofighi, Hamid, Cai, Jianfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
von: Le, Duy-Tho, et al.
Veröffentlicht: (2024)
von: Le, Duy-Tho, et al.
Veröffentlicht: (2024)
DifFUSER: Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation
von: Le, Duy-Tho, et al.
Veröffentlicht: (2024)
von: Le, Duy-Tho, et al.
Veröffentlicht: (2024)
An Empirical Study on How Video-LLMs Answer Video Questions
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025)
How Well Can Vision Language Models See Image Details?
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Any Convex Parametric Shapes
von: Le, Duy-Tho, et al.
Veröffentlicht: (2025)
von: Le, Duy-Tho, et al.
Veröffentlicht: (2025)
ASAP-Textured Gaussians: Enhancing Textured Gaussians with Adaptive Sampling and Anisotropic Parameterization
von: Wei, Meng, et al.
Veröffentlicht: (2025)
von: Wei, Meng, et al.
Veröffentlicht: (2025)
Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering
von: Wei, Meng, et al.
Veröffentlicht: (2024)
von: Wei, Meng, et al.
Veröffentlicht: (2024)
JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2024)
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2024)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
Open-Vocabulary Scene Text Recognition via Pseudo-Image Labeling and Margin Loss
von: Ren, Xuhua, et al.
Veröffentlicht: (2024)
von: Ren, Xuhua, et al.
Veröffentlicht: (2024)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
von: Li, Bin, et al.
Veröffentlicht: (2022)
von: Li, Bin, et al.
Veröffentlicht: (2022)
Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis
von: Sun, Hongyu, et al.
Veröffentlicht: (2025)
von: Sun, Hongyu, et al.
Veröffentlicht: (2025)
APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
MPT: Motion Prompt Tuning for Micro-Expression Recognition
von: Liu, Jiateng, et al.
Veröffentlicht: (2025)
von: Liu, Jiateng, et al.
Veröffentlicht: (2025)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
von: Ehsanpour, Mahsa, et al.
Veröffentlicht: (2024)
von: Ehsanpour, Mahsa, et al.
Veröffentlicht: (2024)
ALLVB: All-in-One Long Video Understanding Benchmark
von: Tan, Xichen, et al.
Veröffentlicht: (2025)
von: Tan, Xichen, et al.
Veröffentlicht: (2025)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
von: Li, Yili, et al.
Veröffentlicht: (2025)
von: Li, Yili, et al.
Veröffentlicht: (2025)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
von: Cheng, Dingxin, et al.
Veröffentlicht: (2024)
von: Cheng, Dingxin, et al.
Veröffentlicht: (2024)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2025)
von: Jahangard, Simindokht, et al.
Veröffentlicht: (2025)
Memory Consolidation Enables Long-Context Video Understanding
von: Balažević, Ivana, et al.
Veröffentlicht: (2024)
von: Balažević, Ivana, et al.
Veröffentlicht: (2024)
Video Panels for Long Video Understanding
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Hallucination Mitigation Prompts Long-term Video Understanding
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
von: Le, Duy-Tho, et al.
Veröffentlicht: (2024) -
DifFUSER: Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation
von: Le, Duy-Tho, et al.
Veröffentlicht: (2024) -
An Empirical Study on How Video-LLMs Answer Video Questions
von: Gou, Chenhui, et al.
Veröffentlicht: (2025) -
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025) -
How Well Can Vision Language Models See Image Details?
von: Gou, Chenhui, et al.
Veröffentlicht: (2024)