VITED: Video Temporal Evidence Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Yujie, Song, Yale, Wang, William, Torresani, Lorenzo, Nagarajan, Tushar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Step Differences in Instructional Video
von: Nagarajan, Tushar, et al.
Veröffentlicht: (2024)
von: Nagarajan, Tushar, et al.
Veröffentlicht: (2024)
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
Video ReCap: Recursive Captioning of Hour-Long Videos
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features
von: Ji, Shihao, et al.
Veröffentlicht: (2025)
von: Ji, Shihao, et al.
Veröffentlicht: (2025)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
von: He, Xuehai, et al.
Veröffentlicht: (2024)
von: He, Xuehai, et al.
Veröffentlicht: (2024)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
von: Li, Shicheng, et al.
Veröffentlicht: (2023)
von: Li, Shicheng, et al.
Veröffentlicht: (2023)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
von: Poppi, Tobia, et al.
Veröffentlicht: (2026)
von: Poppi, Tobia, et al.
Veröffentlicht: (2026)
Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
Spatio-Temporal Side Tuning Pre-trained Foundation Models for Video-based Pedestrian Attribute Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
Semantic Compositions Enhance Vision-Language Contrastive Learning
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
von: Li, Bin, et al.
Veröffentlicht: (2022)
von: Li, Bin, et al.
Veröffentlicht: (2022)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
von: Morini, Marco, et al.
Veröffentlicht: (2026)
von: Morini, Marco, et al.
Veröffentlicht: (2026)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
von: Du, Yang, et al.
Veröffentlicht: (2025)
von: Du, Yang, et al.
Veröffentlicht: (2025)
DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs
von: Chiu, Bo-Cheng, et al.
Veröffentlicht: (2025)
von: Chiu, Bo-Cheng, et al.
Veröffentlicht: (2025)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
Scaling RL to Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2025)
von: Chen, Yukang, et al.
Veröffentlicht: (2025)
A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency
von: Long, Do Xuan, et al.
Veröffentlicht: (2026)
von: Long, Do Xuan, et al.
Veröffentlicht: (2026)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
Enhancing Temporal Modeling of Video LLMs via Time Gating
von: Hu, Zi-Yuan, et al.
Veröffentlicht: (2024)
von: Hu, Zi-Yuan, et al.
Veröffentlicht: (2024)
Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics
von: Zhao, Yinjie, et al.
Veröffentlicht: (2025)
von: Zhao, Yinjie, et al.
Veröffentlicht: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance
von: Verghese, Mrinal, et al.
Veröffentlicht: (2024)
von: Verghese, Mrinal, et al.
Veröffentlicht: (2024)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Step Differences in Instructional Video
von: Nagarajan, Tushar, et al.
Veröffentlicht: (2024) -
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024) -
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025) -
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025) -
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)