MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914971674214400 |
|---|---|
| author | Zhu, Minghao Wang, Zhengpu Hu, Mengxian Dang, Ronghao Lin, Xiao Zhou, Xun Liu, Chengju Chen, Qijun |
| author_facet | Zhu, Minghao Wang, Zhengpu Hu, Mengxian Dang, Ronghao Lin, Xiao Zhou, Xun Liu, Chengju Chen, Qijun |
| contents | Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules are added to capture the temporal information. However, zero-shot generalization diminishes with the increase in the number of specialized parameters, making existing works a trade-off between zero-shot and close-set performance. In this paper, we present MoTE, a novel framework that enables generalization and specialization to be balanced in one unified model. Our approach tunes a mixture of temporal experts to learn multiple task views with various degrees of data fitting. To maximally preserve the knowledge of each expert, we propose \emph{Weight Merging Regularization}, which regularizes the merging process of experts in weight space. Additionally with temporal feature modulation to regularize the contribution of temporal feature during test. We achieve a sound balance between zero-shot and close-set video recognition tasks and obtain state-of-the-art or competitive results on various datasets, including Kinetics-400 \& 600, UCF, and HMDB. Code is available at \url{https://github.com/ZMHH-H/MoTE}. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_10589 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer Zhu, Minghao Wang, Zhengpu Hu, Mengxian Dang, Ronghao Lin, Xiao Zhou, Xun Liu, Chengju Chen, Qijun Computer Vision and Pattern Recognition Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules are added to capture the temporal information. However, zero-shot generalization diminishes with the increase in the number of specialized parameters, making existing works a trade-off between zero-shot and close-set performance. In this paper, we present MoTE, a novel framework that enables generalization and specialization to be balanced in one unified model. Our approach tunes a mixture of temporal experts to learn multiple task views with various degrees of data fitting. To maximally preserve the knowledge of each expert, we propose \emph{Weight Merging Regularization}, which regularizes the merging process of experts in weight space. Additionally with temporal feature modulation to regularize the contribution of temporal feature during test. We achieve a sound balance between zero-shot and close-set video recognition tasks and obtain state-of-the-art or competitive results on various datasets, including Kinetics-400 \& 600, UCF, and HMDB. Code is available at \url{https://github.com/ZMHH-H/MoTE}. |
| title | MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.10589 |