LazyVLM: Neuro-Symbolic Approach to Video Analytics

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jian, Xiangru, Pang, Wei, Dong, Zhengyuan, Zhang, Chao, Özsu, M. Tamer
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910970797883392
author Jian, Xiangru
Pang, Wei
Dong, Zhengyuan
Zhang, Chao
Özsu, M. Tamer
author_facet Jian, Xiangru
Pang, Wei
Dong, Zhengyuan
Zhang, Chao
Özsu, M. Tamer
contents Current video analytics approaches face a fundamental trade-off between flexibility and efficiency. End-to-end Vision Language Models (VLMs) often struggle with long-context processing and incur high computational costs, while neural-symbolic methods depend heavily on manual labeling and rigid rule design. In this paper, we introduce LazyVLM, a neuro-symbolic video analytics system that provides a user-friendly query interface similar to VLMs, while addressing their scalability limitation. LazyVLM enables users to effortlessly drop in video data and specify complex multi-frame video queries using a semi-structured text interface for video analytics. To address the scalability limitations of VLMs, LazyVLM decomposes multi-frame video queries into fine-grained operations and offloads the bulk of the processing to efficient relational query execution and vector similarity search. We demonstrate that LazyVLM provides a robust, efficient, and user-friendly solution for querying open-domain video data at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21459
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LazyVLM: Neuro-Symbolic Approach to Video Analytics
Jian, Xiangru
Pang, Wei
Dong, Zhengyuan
Zhang, Chao
Özsu, M. Tamer
Databases
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
Current video analytics approaches face a fundamental trade-off between flexibility and efficiency. End-to-end Vision Language Models (VLMs) often struggle with long-context processing and incur high computational costs, while neural-symbolic methods depend heavily on manual labeling and rigid rule design. In this paper, we introduce LazyVLM, a neuro-symbolic video analytics system that provides a user-friendly query interface similar to VLMs, while addressing their scalability limitation. LazyVLM enables users to effortlessly drop in video data and specify complex multi-frame video queries using a semi-structured text interface for video analytics. To address the scalability limitations of VLMs, LazyVLM decomposes multi-frame video queries into fine-grained operations and offloads the bulk of the processing to efficient relational query execution and vector similarity search. We demonstrate that LazyVLM provides a robust, efficient, and user-friendly solution for querying open-domain video data at scale.
title LazyVLM: Neuro-Symbolic Approach to Video Analytics
topic Databases
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
url https://arxiv.org/abs/2505.21459