Before Smelling the Video: A Two-Stage Pipeline for Interpretable Video-to-Scent Plans

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Kaicheng, Shao, Kevin Zhongyang, Chen, Ruiqi, Makhsous, Sep, Wilson, Denise
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915757609189376
author Wang, Kaicheng
Shao, Kevin Zhongyang
Chen, Ruiqi
Makhsous, Sep
Wilson, Denise
author_facet Wang, Kaicheng
Shao, Kevin Zhongyang
Chen, Ruiqi
Makhsous, Sep
Wilson, Denise
contents Olfactory cues can enhance immersion in interactive media, yet smell remains rare because it is difficult to author and synchronize with dynamic video. Prior olfactory interfaces rely on designer triggers and fixed event-to-odor mappings that do not scale to unconstrained content. This work examines whether semantic planning for smell is intelligible to people before physical scent delivery. We present a video-to-scent planning pipeline that separates visual semantic extraction using a vision-language model from semantic-to-olfactory inference using a large language model. Two survey studies compare system-generated scent plans with over-inclusive and naive baselines. Results show consistent preference for plans that prioritize perceptually salient cues and align scent changes with visible actions, supporting semantic planning as a foundation for future olfactory media systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19203
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Before Smelling the Video: A Two-Stage Pipeline for Interpretable Video-to-Scent Plans
Wang, Kaicheng
Shao, Kevin Zhongyang
Chen, Ruiqi
Makhsous, Sep
Wilson, Denise
Human-Computer Interaction
H.5.2
Olfactory cues can enhance immersion in interactive media, yet smell remains rare because it is difficult to author and synchronize with dynamic video. Prior olfactory interfaces rely on designer triggers and fixed event-to-odor mappings that do not scale to unconstrained content. This work examines whether semantic planning for smell is intelligible to people before physical scent delivery. We present a video-to-scent planning pipeline that separates visual semantic extraction using a vision-language model from semantic-to-olfactory inference using a large language model. Two survey studies compare system-generated scent plans with over-inclusive and naive baselines. Results show consistent preference for plans that prioritize perceptually salient cues and align scent changes with visible actions, supporting semantic planning as a foundation for future olfactory media systems.
title Before Smelling the Video: A Two-Stage Pipeline for Interpretable Video-to-Scent Plans
topic Human-Computer Interaction
H.5.2
url https://arxiv.org/abs/2601.19203