Chain-of-Anomaly Thoughts with Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Domingos, Pedro, Pereira, João, Lopes, Vasco, Neves, João, Semedo, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909974741909504
author Domingos, Pedro
Pereira, João
Lopes, Vasco
Neves, João
Semedo, David
author_facet Domingos, Pedro
Pereira, João
Lopes, Vasco
Neves, João
Semedo, David
contents Automated video surveillance with Large Vision-Language Models is limited by their inherent bias towards normality, often failing to detect crimes. While Chain-of-Thought reasoning strategies show significant potential for improving performance in language tasks, the lack of inductive anomaly biases in their reasoning further steers the models towards normal interpretations. To address this, we propose Chain-of-Anomaly-Thoughts (CoAT), a multi-agent reasoning framework that introduces inductive criminal bias in the reasoning process through a final, anomaly-focused classification layer. Our method significantly improves Anomaly Detection, boosting F1-score by 11.8 p.p. on challenging low-resolution footage and Anomaly Classification by 3.78 p.p. in high-resolution videos.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Anomaly Thoughts with Large Vision-Language Models
Domingos, Pedro
Pereira, João
Lopes, Vasco
Neves, João
Semedo, David
Computer Vision and Pattern Recognition
Multiagent Systems
Automated video surveillance with Large Vision-Language Models is limited by their inherent bias towards normality, often failing to detect crimes. While Chain-of-Thought reasoning strategies show significant potential for improving performance in language tasks, the lack of inductive anomaly biases in their reasoning further steers the models towards normal interpretations. To address this, we propose Chain-of-Anomaly-Thoughts (CoAT), a multi-agent reasoning framework that introduces inductive criminal bias in the reasoning process through a final, anomaly-focused classification layer. Our method significantly improves Anomaly Detection, boosting F1-score by 11.8 p.p. on challenging low-resolution footage and Anomaly Classification by 3.78 p.p. in high-resolution videos.
title Chain-of-Anomaly Thoughts with Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Multiagent Systems
url https://arxiv.org/abs/2512.20417