Saved in:
Bibliographic Details
Main Authors: Ghosh, Rahul, Chaudhury, Baishali, Das, Hari Prasanna, Ashok, Meghana, Razkenari, Ryan, Chen, Long, Hong, Sungmin, Liu, Chun-Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.00171
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912974773420032
author Ghosh, Rahul
Chaudhury, Baishali
Das, Hari Prasanna
Ashok, Meghana
Razkenari, Ryan
Chen, Long
Hong, Sungmin
Liu, Chun-Hao
author_facet Ghosh, Rahul
Chaudhury, Baishali
Das, Hari Prasanna
Ashok, Meghana
Razkenari, Ryan
Chen, Long
Hong, Sungmin
Liu, Chun-Hao
contents Visual compliance verification is a critical yet underexplored problem in computer vision, especially in domains such as media, entertainment, and advertising where content must adhere to complex and evolving policy rules. Existing methods often rely on task-specific deep learning models trained on manually labeled datasets, which are costly to build and limited in generalizability. While recent Multimodal Large Language Models (MLLMs) offer broad real-world knowledge and policy understanding, they struggle to reason over fine-grained visual details and apply structured compliance rules effectively on their own. In this paper, we propose CompAgent, the first agentic framework for visual compliance verification. CompAgent augments MLLMs with a suite of visual tools-such as object detectors, face analyzers, NSFW detectors, and captioning models-and introduces a planning agent that dynamically selects appropriate tools based on the compliance policy. A compliance verification agent then integrates image, tool outputs, and policy context to perform multimodal reasoning. Experiments on public benchmarks show that CompAgent outperforms specialized classifiers, direct MLLM prompting, and curated routing baselines, achieving up to 76% F1 score and a 10% improvement over the state-of-the-art on the UnsafeBench dataset. Our results demonstrate the effectiveness of agentic planning and robust tool-augmented reasoning for scalable, accurate, and adaptable visual compliance verification.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00171
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CompAgent: An Agentic Framework for Visual Compliance Verification
Ghosh, Rahul
Chaudhury, Baishali
Das, Hari Prasanna
Ashok, Meghana
Razkenari, Ryan
Chen, Long
Hong, Sungmin
Liu, Chun-Hao
Computer Vision and Pattern Recognition
Visual compliance verification is a critical yet underexplored problem in computer vision, especially in domains such as media, entertainment, and advertising where content must adhere to complex and evolving policy rules. Existing methods often rely on task-specific deep learning models trained on manually labeled datasets, which are costly to build and limited in generalizability. While recent Multimodal Large Language Models (MLLMs) offer broad real-world knowledge and policy understanding, they struggle to reason over fine-grained visual details and apply structured compliance rules effectively on their own. In this paper, we propose CompAgent, the first agentic framework for visual compliance verification. CompAgent augments MLLMs with a suite of visual tools-such as object detectors, face analyzers, NSFW detectors, and captioning models-and introduces a planning agent that dynamically selects appropriate tools based on the compliance policy. A compliance verification agent then integrates image, tool outputs, and policy context to perform multimodal reasoning. Experiments on public benchmarks show that CompAgent outperforms specialized classifiers, direct MLLM prompting, and curated routing baselines, achieving up to 76% F1 score and a 10% improvement over the state-of-the-art on the UnsafeBench dataset. Our results demonstrate the effectiveness of agentic planning and robust tool-augmented reasoning for scalable, accurate, and adaptable visual compliance verification.
title CompAgent: An Agentic Framework for Visual Compliance Verification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.00171