Saved in:
Bibliographic Details
Main Authors: Kim, Minsun, Lee, Dawon, Noh, Junyong
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.26173
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • On general video-sharing platforms like YouTube, comments are displayed independently of video playback. As viewers often read comments while watching a video, they may encounter ones referring to moments unrelated to the current scene, which can reveal spoilers and disrupt immersion. To address this problem, we present ComVi, a novel system that displays comments at contextually relevant moments, enabling viewers to see time-synchronized comments and video content together. We first map all comments to relevant video timestamps by computing audio-visual correlation, then construct the comment sequence through an optimization that considers temporal relevance, popularity (number of likes), and display duration for comfortable reading. In a user study, ComVi provided a significantly more engaging experience than conventional video interfaces (i.e., YouTube and Danmaku), with 71.9% of participants selecting ComVi as their most preferred interface.