Embedding-based Retrieval in Multimodal Content Moderation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Hanzhong, Shi, Jinghao, Shen, Xiang, Wang, Zixuan, Wen, Vera, Mehrani, Ardalan, Chen, Zhiqian, Wu, Yifan, Zhang, Zhixin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909671596490752
author Liang, Hanzhong
Shi, Jinghao
Shen, Xiang
Wang, Zixuan
Wen, Vera
Mehrani, Ardalan
Chen, Zhiqian
Wu, Yifan
Zhang, Zhixin
author_facet Liang, Hanzhong
Shi, Jinghao
Shen, Xiang
Wang, Zixuan
Wen, Vera
Mehrani, Ardalan
Chen, Zhiqian
Wu, Yifan
Zhang, Zhixin
contents Video understanding plays a fundamental role for content moderation on short video platforms, enabling the detection of inappropriate content. While classification remains the dominant approach for content moderation, it often struggles in scenarios requiring rapid and cost-efficient responses, such as trend adaptation and urgent escalations. To address this issue, we introduce an Embedding-Based Retrieval (EBR) method designed to complement traditional classification approaches. We first leverage a Supervised Contrastive Learning (SCL) framework to train a suite of foundation embedding models, including both single-modal and multi-modal architectures. Our models demonstrate superior performance over established contrastive learning methods such as CLIP and MoCo. Building on these embedding models, we design and implement the embedding-based retrieval system that integrates embedding generation and video retrieval to enable efficient and effective trend handling. Comprehensive offline experiments on 25 diverse emerging trends show that EBR improves ROC-AUC from 0.85 to 0.99 and PR-AUC from 0.35 to 0.95. Further online experiments reveal that EBR increases action rates by 10.32% and reduces operational costs by over 80%, while also enhancing interpretability and flexibility compared to classification-based solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01066
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Embedding-based Retrieval in Multimodal Content Moderation
Liang, Hanzhong
Shi, Jinghao
Shen, Xiang
Wang, Zixuan
Wen, Vera
Mehrani, Ardalan
Chen, Zhiqian
Wu, Yifan
Zhang, Zhixin
Information Retrieval
Computer Vision and Pattern Recognition
Machine Learning
Video understanding plays a fundamental role for content moderation on short video platforms, enabling the detection of inappropriate content. While classification remains the dominant approach for content moderation, it often struggles in scenarios requiring rapid and cost-efficient responses, such as trend adaptation and urgent escalations. To address this issue, we introduce an Embedding-Based Retrieval (EBR) method designed to complement traditional classification approaches. We first leverage a Supervised Contrastive Learning (SCL) framework to train a suite of foundation embedding models, including both single-modal and multi-modal architectures. Our models demonstrate superior performance over established contrastive learning methods such as CLIP and MoCo. Building on these embedding models, we design and implement the embedding-based retrieval system that integrates embedding generation and video retrieval to enable efficient and effective trend handling. Comprehensive offline experiments on 25 diverse emerging trends show that EBR improves ROC-AUC from 0.85 to 0.99 and PR-AUC from 0.35 to 0.95. Further online experiments reveal that EBR increases action rates by 10.32% and reduces operational costs by over 80%, while also enhancing interpretability and flexibility compared to classification-based solutions.
title Embedding-based Retrieval in Multimodal Content Moderation
topic Information Retrieval
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.01066