Large Language Models for Video Surveillance Applications
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | De Silva, Ulindu, Fernando, Leon, Lik, Billy Lau Pik, Koh, Zann, Joyce, Sam Conrad, Yuen, Belinda, Yuen, Chau |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Scalable Decentralized Reinforcement Learning Framework for UAV Target Localization Using Recurrent PPO
par: Fernando, Leon, et autres
Publié: (2024)
par: Fernando, Leon, et autres
Publié: (2024)
Video Summarisation with Incident and Context Information using Generative AI
par: De Silva, Ulindu, et autres
Publié: (2025)
par: De Silva, Ulindu, et autres
Publié: (2025)
MEF-Explore: Communication-Constrained Multi-Robot Entropy-Field-Based Exploration
par: Pongsirijinda, Khattiya, et autres
Publié: (2025)
par: Pongsirijinda, Khattiya, et autres
Publié: (2025)
Selective State Space Memory for Large Vision-Language Models
par: Ng, Chee, et autres
Publié: (2024)
par: Ng, Chee, et autres
Publié: (2024)
Distributed multi-robot potential-field-based exploration with submap-based mapping and noise-augmented strategy
par: Pongsirijinda, Khattiya, et autres
Publié: (2024)
par: Pongsirijinda, Khattiya, et autres
Publié: (2024)
ENTED: Enhanced Neural Texture Extraction and Distribution for Reference-based Blind Face Restoration
par: Lau, Yuen-Fui, et autres
Publié: (2024)
par: Lau, Yuen-Fui, et autres
Publié: (2024)
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
par: Liu, Bo, et autres
Publié: (2025)
par: Liu, Bo, et autres
Publié: (2025)
Leveraging RGB-D Data with Cross-Modal Context Mining for Glass Surface Detection
par: Lin, Jiaying, et autres
Publié: (2022)
par: Lin, Jiaying, et autres
Publié: (2022)
A Benchmark for Crime Surveillance Video Analysis with Large Models
par: Chen, Haoran, et autres
Publié: (2025)
par: Chen, Haoran, et autres
Publié: (2025)
Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation
par: De Silva, Ulindu, et autres
Publié: (2025)
par: De Silva, Ulindu, et autres
Publié: (2025)
Entropy Guided Dynamic Patch Segmentation for Time Series Transformers
par: Abeywickrama, Sachith, et autres
Publié: (2025)
par: Abeywickrama, Sachith, et autres
Publié: (2025)
Detection of Object Throwing Behavior in Surveillance Videos
par: Kersten, Ivo P. C., et autres
Publié: (2024)
par: Kersten, Ivo P. C., et autres
Publié: (2024)
Personalized Large Vision-Language Models
par: Pham, Chau, et autres
Publié: (2024)
par: Pham, Chau, et autres
Publié: (2024)
A Tutorial on Explainable Image Classification for Dementia Stages Using Convolutional Neural Network and Gradient-weighted Class Activation Mapping
par: Yuen, Kevin Kam Fung
Publié: (2024)
par: Yuen, Kevin Kam Fung
Publié: (2024)
Large-scale Self-supervised Video Foundation Model for Intelligent Surgery
par: Yang, Shu, et autres
Publié: (2025)
par: Yang, Shu, et autres
Publié: (2025)
Real-Time AI-Driven People Tracking and Counting Using Overhead Cameras
par: Ahamed, Ishrath, et autres
Publié: (2024)
par: Ahamed, Ishrath, et autres
Publié: (2024)
Exploring Causes and Mitigation of Hallucinations in Large Vision Language Models
par: Sun, Yaqi, et autres
Publié: (2025)
par: Sun, Yaqi, et autres
Publié: (2025)
SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
par: Rabasseda, Oriol, et autres
Publié: (2026)
par: Rabasseda, Oriol, et autres
Publié: (2026)
Evaluation of Vision-LLMs in Surveillance Video
par: Benschop, Pascal, et autres
Publié: (2025)
par: Benschop, Pascal, et autres
Publié: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
par: Zhang, Pingping, et autres
Publié: (2024)
par: Zhang, Pingping, et autres
Publié: (2024)
ProDisc-VAD: An Efficient System for Weakly-Supervised Anomaly Detection in Video Surveillance Applications
par: Zhu, Tao, et autres
Publié: (2025)
par: Zhu, Tao, et autres
Publié: (2025)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
par: Sun, Haoran, et autres
Publié: (2025)
par: Sun, Haoran, et autres
Publié: (2025)
MTFL: Multi-Timescale Feature Learning for Weakly-Supervised Anomaly Detection in Surveillance Videos
par: Zhang, Yiling, et autres
Publié: (2024)
par: Zhang, Yiling, et autres
Publié: (2024)
Video Summarization with Large Language Models
par: Lee, Min Jung, et autres
Publié: (2025)
par: Lee, Min Jung, et autres
Publié: (2025)
C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning
par: Ye, Guanting, et autres
Publié: (2026)
par: Ye, Guanting, et autres
Publié: (2026)
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024)
par: Chen, Joya, et autres
Publié: (2024)
Recognition of Abnormal Events in Surveillance Videos using Weakly Supervised Dual-Encoder Models
par: Tsfaty, Noam, et autres
Publié: (2025)
par: Tsfaty, Noam, et autres
Publié: (2025)
Video-based Vehicle Surveillance in the Wild: License Plate, Make, and Model Recognition with Self Reflective Vision-Language Models
par: Parsa, Pouya, et autres
Publié: (2025)
par: Parsa, Pouya, et autres
Publié: (2025)
A Flying Bird Object Detection Method for Surveillance Video
par: Sun, Ziwei, et autres
Publié: (2024)
par: Sun, Ziwei, et autres
Publié: (2024)
AssistPDA: An Online Video Surveillance Assistant for Video Anomaly Prediction, Detection, and Analysis
par: Yang, Zhiwei, et autres
Publié: (2025)
par: Yang, Zhiwei, et autres
Publié: (2025)
Improved Image-based Pose Regressor Models for Underwater Environments
par: Peng, Luyuan, et autres
Publié: (2024)
par: Peng, Luyuan, et autres
Publié: (2024)
Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
par: Barbara, Mario, et autres
Publié: (2025)
par: Barbara, Mario, et autres
Publié: (2025)
Zero-Shot Action Recognition in Surveillance Videos
par: Pereira, Joao, et autres
Publié: (2024)
par: Pereira, Joao, et autres
Publié: (2024)
Dynamic Against Dynamic: An Open-set Self-learning Framework
par: Yang, Haifeng, et autres
Publié: (2024)
par: Yang, Haifeng, et autres
Publié: (2024)
On the Consistency of Video Large Language Models in Temporal Comprehension
par: Jung, Minjoon, et autres
Publié: (2024)
par: Jung, Minjoon, et autres
Publié: (2024)
Streaming Long Video Understanding with Large Language Models
par: Qian, Rui, et autres
Publié: (2024)
par: Qian, Rui, et autres
Publié: (2024)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
par: Cao, Meng, et autres
Publié: (2025)
par: Cao, Meng, et autres
Publié: (2025)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
par: Wang, Yuxuan, et autres
Publié: (2024)
par: Wang, Yuxuan, et autres
Publié: (2024)
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
par: He, Zhihao, et autres
Publié: (2025)
par: He, Zhihao, et autres
Publié: (2025)
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
par: De Nadai, Marco, et autres
Publié: (2025)
par: De Nadai, Marco, et autres
Publié: (2025)
Documents similaires
-
A Scalable Decentralized Reinforcement Learning Framework for UAV Target Localization Using Recurrent PPO
par: Fernando, Leon, et autres
Publié: (2024) -
Video Summarisation with Incident and Context Information using Generative AI
par: De Silva, Ulindu, et autres
Publié: (2025) -
MEF-Explore: Communication-Constrained Multi-Robot Entropy-Field-Based Exploration
par: Pongsirijinda, Khattiya, et autres
Publié: (2025) -
Selective State Space Memory for Large Vision-Language Models
par: Ng, Chee, et autres
Publié: (2024) -
Distributed multi-robot potential-field-based exploration with submap-based mapping and noise-augmented strategy
par: Pongsirijinda, Khattiya, et autres
Publié: (2024)