DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nguyen, Thong, Wu, Xiaobao, Dong, Xinshuai, Nguyen, Cong-Duy, Ng, See-Kiong, Tuan, Luu Anh
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913114333642752
author Nguyen, Thong
Wu, Xiaobao
Dong, Xinshuai
Nguyen, Cong-Duy
Ng, See-Kiong
Tuan, Luu Anh
author_facet Nguyen, Thong
Wu, Xiaobao
Dong, Xinshuai
Nguyen, Cong-Duy
Ng, See-Kiong
Tuan, Luu Anh
contents Temporal Language Grounding seeks to localize video moments that semantically correspond to a natural language query. Recent advances employ the attention mechanism to learn the relations between video moments and the text query. However, naive attention might not be able to appropriately capture such relations, resulting in ineffective distributions where target video moments are difficult to separate from the remaining ones. To resolve the issue, we propose an energy-based model framework to explicitly learn moment-query distributions. Moreover, we propose DemaFormer, a novel Transformer-based architecture that utilizes exponential moving average with a learnable damping factor to effectively encode moment-query inputs. Comprehensive experiments on four public temporal language grounding datasets showcase the superiority of our methods over the state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02549
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
Nguyen, Thong
Wu, Xiaobao
Dong, Xinshuai
Nguyen, Cong-Duy
Ng, See-Kiong
Tuan, Luu Anh
Computer Vision and Pattern Recognition
Computation and Language
Temporal Language Grounding seeks to localize video moments that semantically correspond to a natural language query. Recent advances employ the attention mechanism to learn the relations between video moments and the text query. However, naive attention might not be able to appropriately capture such relations, resulting in ineffective distributions where target video moments are difficult to separate from the remaining ones. To resolve the issue, we propose an energy-based model framework to explicitly learn moment-query distributions. Moreover, we propose DemaFormer, a novel Transformer-based architecture that utilizes exponential moving average with a learnable damping factor to effectively encode moment-query inputs. Comprehensive experiments on four public temporal language grounding datasets showcase the superiority of our methods over the state-of-the-art baselines.
title DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2312.02549