Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement Trajectories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asano, Hikaru, Yonetani, Ryo, Sekii, Taiki, Ouchi, Hiroki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912035012345856
author Asano, Hikaru
Yonetani, Ryo
Sekii, Taiki
Ouchi, Hiroki
author_facet Asano, Hikaru
Yonetani, Ryo
Sekii, Taiki
Ouchi, Hiroki
contents This paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores. Our work will impact various retail applications that need better customer understanding, such as targeted advertising and inventory management. The key idea is leveraging large language models to synthesize a diverse and realistic collection of contextual captions as well as the corresponding movement trajectories on a store map. Despite learned from fully synthesized data, the captioning model can generalize well to trajectories/captions created by real human subjects. Our systematic evaluation confirmed the effectiveness of the proposed framework over competitive approaches in terms of ROUGE and BERT Score metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12670
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement Trajectories
Asano, Hikaru
Yonetani, Ryo
Sekii, Taiki
Ouchi, Hiroki
Computation and Language
This paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores. Our work will impact various retail applications that need better customer understanding, such as targeted advertising and inventory management. The key idea is leveraging large language models to synthesize a diverse and realistic collection of contextual captions as well as the corresponding movement trajectories on a store map. Despite learned from fully synthesized data, the captioning model can generalize well to trajectories/captions created by real human subjects. Our systematic evaluation confirmed the effectiveness of the proposed framework over competitive approaches in terms of ROUGE and BERT Score metrics.
title Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement Trajectories
topic Computation and Language
url https://arxiv.org/abs/2409.12670