Robustness Evaluation for Video Models with Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Babu, Ashwin Ramesh, Mousavi, Sajad, Gundecha, Vineet, Ghorbanpour, Sahand, Naug, Avisek, Guillen, Antonio, Gutierrez, Ricardo Luna, Sarkar, Soumyendu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918046599217152
author Babu, Ashwin Ramesh
Mousavi, Sajad
Gundecha, Vineet
Ghorbanpour, Sahand
Naug, Avisek
Guillen, Antonio
Gutierrez, Ricardo Luna
Sarkar, Soumyendu
author_facet Babu, Ashwin Ramesh
Mousavi, Sajad
Gundecha, Vineet
Ghorbanpour, Sahand
Naug, Avisek
Guillen, Antonio
Gutierrez, Ricardo Luna
Sarkar, Soumyendu
contents Evaluating the robustness of Video classification models is very challenging, specifically when compared to image-based models. With their increased temporal dimension, there is a significant increase in complexity and computational cost. One of the key challenges is to keep the perturbations to a minimum to induce misclassification. In this work, we propose a multi-agent reinforcement learning approach (spatial and temporal) that cooperatively learns to identify the given video's sensitive spatial and temporal regions. The agents consider temporal coherence in generating fine perturbations, leading to a more effective and visually imperceptible attack. Our method outperforms the state-of-the-art solutions on the Lp metric and the average queries. Our method enables custom distortion types, making the robustness evaluation more relevant to the use case. We extensively evaluate 4 popular models for video action recognition on two popular datasets, HMDB-51 and UCF-101.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05431
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robustness Evaluation for Video Models with Reinforcement Learning
Babu, Ashwin Ramesh
Mousavi, Sajad
Gundecha, Vineet
Ghorbanpour, Sahand
Naug, Avisek
Guillen, Antonio
Gutierrez, Ricardo Luna
Sarkar, Soumyendu
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Evaluating the robustness of Video classification models is very challenging, specifically when compared to image-based models. With their increased temporal dimension, there is a significant increase in complexity and computational cost. One of the key challenges is to keep the perturbations to a minimum to induce misclassification. In this work, we propose a multi-agent reinforcement learning approach (spatial and temporal) that cooperatively learns to identify the given video's sensitive spatial and temporal regions. The agents consider temporal coherence in generating fine perturbations, leading to a more effective and visually imperceptible attack. Our method outperforms the state-of-the-art solutions on the Lp metric and the average queries. Our method enables custom distortion types, making the robustness evaluation more relevant to the use case. We extensively evaluate 4 popular models for video action recognition on two popular datasets, HMDB-51 and UCF-101.
title Robustness Evaluation for Video Models with Reinforcement Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.05431