VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Ruoliu, Wu, Chu, Shan, Caifeng, He, Ran, Fu, Chaoyou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PersonaVLM: Long-Term Personalized Multimodal LLMs
von: Nie, Chang, et al.
Veröffentlicht: (2026)
von: Nie, Chang, et al.
Veröffentlicht: (2026)
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
von: Li, Lijiang, et al.
Veröffentlicht: (2026)
von: Li, Lijiang, et al.
Veröffentlicht: (2026)
QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2025)
von: Luo, Yongdong, et al.
Veröffentlicht: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
von: Du, Henghui, et al.
Veröffentlicht: (2025)
von: Du, Henghui, et al.
Veröffentlicht: (2025)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
von: Yin, Shukang, et al.
Veröffentlicht: (2024)
von: Yin, Shukang, et al.
Veröffentlicht: (2024)
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
von: zhang, Kaixin, et al.
Veröffentlicht: (2026)
von: zhang, Kaixin, et al.
Veröffentlicht: (2026)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts
von: Wang, Songping, et al.
Veröffentlicht: (2026)
von: Wang, Songping, et al.
Veröffentlicht: (2026)
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection
von: Cai, Zhaolin, et al.
Veröffentlicht: (2025)
von: Cai, Zhaolin, et al.
Veröffentlicht: (2025)
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
von: Dong, Shaoqi, et al.
Veröffentlicht: (2025)
von: Dong, Shaoqi, et al.
Veröffentlicht: (2025)
Facial Identity Anonymization via Intrinsic and Extrinsic Attention Distraction
von: Kuang, Zhenzhong, et al.
Veröffentlicht: (2024)
von: Kuang, Zhenzhong, et al.
Veröffentlicht: (2024)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
von: Bai, Purui, et al.
Veröffentlicht: (2026)
von: Bai, Purui, et al.
Veröffentlicht: (2026)
LongVLM: Efficient Long Video Understanding via Large Language Models
von: Weng, Yuetian, et al.
Veröffentlicht: (2024)
von: Weng, Yuetian, et al.
Veröffentlicht: (2024)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
Video-Zero: Self-Evolution Video Understanding
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
von: You, Zeng, et al.
Veröffentlicht: (2024)
von: You, Zeng, et al.
Veröffentlicht: (2024)
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
Video Panels for Long Video Understanding
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
von: Reza, Sakib, et al.
Veröffentlicht: (2025)
von: Reza, Sakib, et al.
Veröffentlicht: (2025)
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
von: He, Haichen, et al.
Veröffentlicht: (2026)
von: He, Haichen, et al.
Veröffentlicht: (2026)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PersonaVLM: Long-Term Personalized Multimodal LLMs
von: Nie, Chang, et al.
Veröffentlicht: (2026) -
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026) -
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
von: Li, Lijiang, et al.
Veröffentlicht: (2026) -
QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2025) -
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)