Preventing Conflicting Gradients in Neural Marked Temporal Point Processes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bosser, Tanguy, Taieb, Souhaib Ben
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912152478023680
author Bosser, Tanguy
Taieb, Souhaib Ben
author_facet Bosser, Tanguy
Taieb, Souhaib Ben
contents Neural Marked Temporal Point Processes (MTPP) are flexible models to capture complex temporal inter-dependencies between labeled events. These models inherently learn two predictive distributions: one for the arrival times of events and another for the types of events, also known as marks. In this study, we demonstrate that learning a MTPP model can be framed as a two-task learning problem, where both tasks share a common set of trainable parameters that are optimized jointly. We show that this often leads to the emergence of conflicting gradients during training, where task-specific gradients are pointing in opposite directions. When such conflicts arise, following the average gradient can be detrimental to the learning of each individual tasks, resulting in overall degraded performance. To overcome this issue, we introduce novel parametrizations for neural MTPP models that allow for separate modeling and training of each task, effectively avoiding the problem of conflicting gradients. Through experiments on multiple real-world event sequence datasets, we demonstrate the benefits of our framework compared to the original model formulations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08590
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Preventing Conflicting Gradients in Neural Marked Temporal Point Processes
Bosser, Tanguy
Taieb, Souhaib Ben
Machine Learning
Neural Marked Temporal Point Processes (MTPP) are flexible models to capture complex temporal inter-dependencies between labeled events. These models inherently learn two predictive distributions: one for the arrival times of events and another for the types of events, also known as marks. In this study, we demonstrate that learning a MTPP model can be framed as a two-task learning problem, where both tasks share a common set of trainable parameters that are optimized jointly. We show that this often leads to the emergence of conflicting gradients during training, where task-specific gradients are pointing in opposite directions. When such conflicts arise, following the average gradient can be detrimental to the learning of each individual tasks, resulting in overall degraded performance. To overcome this issue, we introduce novel parametrizations for neural MTPP models that allow for separate modeling and training of each task, effectively avoiding the problem of conflicting gradients. Through experiments on multiple real-world event sequence datasets, we demonstrate the benefits of our framework compared to the original model formulations.
title Preventing Conflicting Gradients in Neural Marked Temporal Point Processes
topic Machine Learning
url https://arxiv.org/abs/2412.08590