Text-guided Fine-Grained Video Anomaly Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Jihao, Li, Kun, Wang, He, Akşit, Kaan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MA-Bench: Towards Fine-grained Micro-Action Understanding
von: Li, Kun, et al.
Veröffentlicht: (2026)
von: Li, Kun, et al.
Veröffentlicht: (2026)
Editing Physiological Signals in Videos Using Latent Representations
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025)
Learned Display Radiance Fields with Lensless Cameras
von: Chen, Ziyang, et al.
Veröffentlicht: (2025)
von: Chen, Ziyang, et al.
Veröffentlicht: (2025)
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
von: Pereira, João, et al.
Veröffentlicht: (2026)
von: Pereira, João, et al.
Veröffentlicht: (2026)
Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
Learned Single-Pass Multitasking Perceptual Graphics for Immersive Displays
von: Yılmaz, Doğa, et al.
Veröffentlicht: (2024)
von: Yılmaz, Doğa, et al.
Veröffentlicht: (2024)
Performance Analysis of Traditional VQA Models Under Limited Computational Resources
von: Gu, Jihao
Veröffentlicht: (2025)
von: Gu, Jihao
Veröffentlicht: (2025)
Complex-Valued 2D Gaussian Representation for Computer-Generated Holography
von: Zhan, Yicheng, et al.
Veröffentlicht: (2025)
von: Zhan, Yicheng, et al.
Veröffentlicht: (2025)
SpecTrack: Learned Multi-Rotation Tracking via Speckle Imaging
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
von: Peng, Yi-Xing, et al.
Veröffentlicht: (2025)
von: Peng, Yi-Xing, et al.
Veröffentlicht: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
von: Cheng, Zheng, et al.
Veröffentlicht: (2024)
von: Cheng, Zheng, et al.
Veröffentlicht: (2024)
TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2024)
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2024)
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
von: Ponbagavathi, Thinesh Thiyakesan, et al.
Veröffentlicht: (2026)
von: Ponbagavathi, Thinesh Thiyakesan, et al.
Veröffentlicht: (2026)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
von: Zhu, Liyun, et al.
Veröffentlicht: (2025)
von: Zhu, Liyun, et al.
Veröffentlicht: (2025)
FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
von: Pan, Hewen, et al.
Veröffentlicht: (2025)
von: Pan, Hewen, et al.
Veröffentlicht: (2025)
PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
von: Wang, Penghao, et al.
Veröffentlicht: (2025)
von: Wang, Penghao, et al.
Veröffentlicht: (2025)
Complex-Valued Holographic Radiance Fields
von: Zhan, Yicheng, et al.
Veröffentlicht: (2025)
von: Zhan, Yicheng, et al.
Veröffentlicht: (2025)
EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval
von: Chen, Yuhan, et al.
Veröffentlicht: (2026)
von: Chen, Yuhan, et al.
Veröffentlicht: (2026)
VADMamba++: Efficient Video Anomaly Detection via Hybrid Modeling in Grayscale Space
von: Lyu, Jihao, et al.
Veröffentlicht: (2026)
von: Lyu, Jihao, et al.
Veröffentlicht: (2026)
Dissolving Is Amplifying: Towards Fine-Grained Anomaly Detection
von: Shi, Jian, et al.
Veröffentlicht: (2023)
von: Shi, Jian, et al.
Veröffentlicht: (2023)
FineXtrol: Controllable Motion Generation via Fine-Grained Text
von: Shen, Keming, et al.
Veröffentlicht: (2025)
von: Shen, Keming, et al.
Veröffentlicht: (2025)
TextToucher: Fine-Grained Text-to-Touch Generation
von: Tu, Jiahang, et al.
Veröffentlicht: (2024)
von: Tu, Jiahang, et al.
Veröffentlicht: (2024)
VideoChat: Chat-Centric Video Understanding
von: Li, KunChang, et al.
Veröffentlicht: (2023)
von: Li, KunChang, et al.
Veröffentlicht: (2023)
Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
von: Fan, Yuanting, et al.
Veröffentlicht: (2025)
von: Fan, Yuanting, et al.
Veröffentlicht: (2025)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
von: Chen, Houlun, et al.
Veröffentlicht: (2024)
CheXLearner: Text-Guided Fine-Grained Representation Learning for Progression Detection
von: Wang, Yuanzhuo, et al.
Veröffentlicht: (2025)
von: Wang, Yuanzhuo, et al.
Veröffentlicht: (2025)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2024)
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2024)
MotionCharacter: Fine-Grained Motion Controllable Human Video Generation
von: Fang, Haopeng, et al.
Veröffentlicht: (2024)
von: Fang, Haopeng, et al.
Veröffentlicht: (2024)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
von: Chen, Feng, et al.
Veröffentlicht: (2025)
von: Chen, Feng, et al.
Veröffentlicht: (2025)
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
von: wang, Shuo, et al.
Veröffentlicht: (2025)
von: wang, Shuo, et al.
Veröffentlicht: (2025)
Language-guided Open-world Video Anomaly Detection under Weak Supervision
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
Text-Video Multi-Grained Integration for Video Moment Montage
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
von: Zou, Shu, et al.
Veröffentlicht: (2025)
von: Zou, Shu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MA-Bench: Towards Fine-grained Micro-Action Understanding
von: Li, Kun, et al.
Veröffentlicht: (2026) -
Editing Physiological Signals in Videos Using Latent Representations
von: Zhou, Tianwen, et al.
Veröffentlicht: (2025) -
Learned Display Radiance Fields with Lensless Cameras
von: Chen, Ziyang, et al.
Veröffentlicht: (2025) -
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
von: Pereira, João, et al.
Veröffentlicht: (2026) -
Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
von: Gu, Jihao, et al.
Veröffentlicht: (2025)