VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Shibo, Yang, Peipei, Liu, Yangyang, Chen, Yi, Zhu, Han, Zhang, Xuyao, Huang, Linlin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2024)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
von: Zhang, Long, et al.
Veröffentlicht: (2025)
von: Zhang, Long, et al.
Veröffentlicht: (2025)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
InstructHumans: Editing Animated 3D Human Textures with Instructions
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
von: Huang, Dawei, et al.
Veröffentlicht: (2025)
von: Huang, Dawei, et al.
Veröffentlicht: (2025)
Do Joint Audio-Video Generation Models Understand Physics?
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
von: Zhu, Haodong, et al.
Veröffentlicht: (2025)
von: Zhu, Haodong, et al.
Veröffentlicht: (2025)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
von: Qin, You, et al.
Veröffentlicht: (2024)
von: Qin, You, et al.
Veröffentlicht: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Joint Flow And Feature Refinement Using Attention For Video Restoration
von: Merugu, Ranjith, et al.
Veröffentlicht: (2025)
von: Merugu, Ranjith, et al.
Veröffentlicht: (2025)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
von: Zhou, Pengyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Pengyuan, et al.
Veröffentlicht: (2024)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
Face Consistency Benchmark for GenAI Video
von: Podstawski, Michal, et al.
Veröffentlicht: (2025)
von: Podstawski, Michal, et al.
Veröffentlicht: (2025)
TAVGBench: Benchmarking Text to Audible-Video Generation
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
von: Han, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Han, Zhiyuan, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
A Sleep Monitoring System Based on Audio, Video and Depth Information
von: Chen, Lyn Chao-ling, et al.
Veröffentlicht: (2025)
von: Chen, Lyn Chao-ling, et al.
Veröffentlicht: (2025)
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024) -
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
von: Zhu, Jiaqi, et al.
Veröffentlicht: (2024) -
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
von: Chao, Jianghan, et al.
Veröffentlicht: (2025) -
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025) -
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
von: Zhang, Long, et al.
Veröffentlicht: (2025)