LazyVLM: Neuro-Symbolic Approach to Video Analytics
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910970797883392 |
|---|---|
| author | Jian, Xiangru Pang, Wei Dong, Zhengyuan Zhang, Chao Özsu, M. Tamer |
| author_facet | Jian, Xiangru Pang, Wei Dong, Zhengyuan Zhang, Chao Özsu, M. Tamer |
| contents | Current video analytics approaches face a fundamental trade-off between flexibility and efficiency. End-to-end Vision Language Models (VLMs) often struggle with long-context processing and incur high computational costs, while neural-symbolic methods depend heavily on manual labeling and rigid rule design. In this paper, we introduce LazyVLM, a neuro-symbolic video analytics system that provides a user-friendly query interface similar to VLMs, while addressing their scalability limitation. LazyVLM enables users to effortlessly drop in video data and specify complex multi-frame video queries using a semi-structured text interface for video analytics. To address the scalability limitations of VLMs, LazyVLM decomposes multi-frame video queries into fine-grained operations and offloads the bulk of the processing to efficient relational query execution and vector similarity search. We demonstrate that LazyVLM provides a robust, efficient, and user-friendly solution for querying open-domain video data at scale. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_21459 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LazyVLM: Neuro-Symbolic Approach to Video Analytics Jian, Xiangru Pang, Wei Dong, Zhengyuan Zhang, Chao Özsu, M. Tamer Databases Artificial Intelligence Computer Vision and Pattern Recognition Information Retrieval Multimedia Current video analytics approaches face a fundamental trade-off between flexibility and efficiency. End-to-end Vision Language Models (VLMs) often struggle with long-context processing and incur high computational costs, while neural-symbolic methods depend heavily on manual labeling and rigid rule design. In this paper, we introduce LazyVLM, a neuro-symbolic video analytics system that provides a user-friendly query interface similar to VLMs, while addressing their scalability limitation. LazyVLM enables users to effortlessly drop in video data and specify complex multi-frame video queries using a semi-structured text interface for video analytics. To address the scalability limitations of VLMs, LazyVLM decomposes multi-frame video queries into fine-grained operations and offloads the bulk of the processing to efficient relational query execution and vector similarity search. We demonstrate that LazyVLM provides a robust, efficient, and user-friendly solution for querying open-domain video data at scale. |
| title | LazyVLM: Neuro-Symbolic Approach to Video Analytics |
| topic | Databases Artificial Intelligence Computer Vision and Pattern Recognition Information Retrieval Multimedia |
| url | https://arxiv.org/abs/2505.21459 |