Robustness Evaluation for Video Models with Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918046599217152 |
|---|---|
| author | Babu, Ashwin Ramesh Mousavi, Sajad Gundecha, Vineet Ghorbanpour, Sahand Naug, Avisek Guillen, Antonio Gutierrez, Ricardo Luna Sarkar, Soumyendu |
| author_facet | Babu, Ashwin Ramesh Mousavi, Sajad Gundecha, Vineet Ghorbanpour, Sahand Naug, Avisek Guillen, Antonio Gutierrez, Ricardo Luna Sarkar, Soumyendu |
| contents | Evaluating the robustness of Video classification models is very challenging, specifically when compared to image-based models. With their increased temporal dimension, there is a significant increase in complexity and computational cost. One of the key challenges is to keep the perturbations to a minimum to induce misclassification. In this work, we propose a multi-agent reinforcement learning approach (spatial and temporal) that cooperatively learns to identify the given video's sensitive spatial and temporal regions. The agents consider temporal coherence in generating fine perturbations, leading to a more effective and visually imperceptible attack. Our method outperforms the state-of-the-art solutions on the Lp metric and the average queries. Our method enables custom distortion types, making the robustness evaluation more relevant to the use case. We extensively evaluate 4 popular models for video action recognition on two popular datasets, HMDB-51 and UCF-101. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_05431 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Robustness Evaluation for Video Models with Reinforcement Learning Babu, Ashwin Ramesh Mousavi, Sajad Gundecha, Vineet Ghorbanpour, Sahand Naug, Avisek Guillen, Antonio Gutierrez, Ricardo Luna Sarkar, Soumyendu Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Evaluating the robustness of Video classification models is very challenging, specifically when compared to image-based models. With their increased temporal dimension, there is a significant increase in complexity and computational cost. One of the key challenges is to keep the perturbations to a minimum to induce misclassification. In this work, we propose a multi-agent reinforcement learning approach (spatial and temporal) that cooperatively learns to identify the given video's sensitive spatial and temporal regions. The agents consider temporal coherence in generating fine perturbations, leading to a more effective and visually imperceptible attack. Our method outperforms the state-of-the-art solutions on the Lp metric and the average queries. Our method enables custom distortion types, making the robustness evaluation more relevant to the use case. We extensively evaluate 4 popular models for video action recognition on two popular datasets, HMDB-51 and UCF-101. |
| title | Robustness Evaluation for Video Models with Reinforcement Learning |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2506.05431 |