Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jäckl, Bastian, Kruchina, Jiří, Joos, Lucas, Keim, Daniel A., Peška, Ladislav, Lokoč, Jakub
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914076355985408
author Jäckl, Bastian
Kruchina, Jiří
Joos, Lucas
Keim, Daniel A.
Peška, Ladislav
Lokoč, Jakub
author_facet Jäckl, Bastian
Kruchina, Jiří
Joos, Lucas
Keim, Daniel A.
Peška, Ladislav
Lokoč, Jakub
contents Multimodal deep-learning models power interactive video retrieval by ranking keyframes in response to textual queries. Despite these advances, users must still browse ranked candidates manually to locate a target. Keyframe arrangement within the search grid highly affects browsing effectiveness and user efficiency, yet remains underexplored. We report a study with 49 participants evaluating seven keyframe layouts for the Visual Known-Item Search task. Beyond efficiency and accuracy, we relate browsing phenomena, such as overlooks, to layout characteristics. Our results show that a video-grouped layout is the most efficient, while a four-column, rank-preserving grid achieves the highest accuracy. Sorted grids reveal potentials and trade-offs, enabling rapid scanning of uninteresting regions but down-ranking relevant targets to less prominent positions, delaying first arrival times and increasing overlooks. These findings motivate hybrid designs that preserve positions of top-ranked items while sorting or grouping the remainder, and offer guidance for searching in grids beyond video retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04396
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections
Jäckl, Bastian
Kruchina, Jiří
Joos, Lucas
Keim, Daniel A.
Peška, Ladislav
Lokoč, Jakub
Multimedia
Information Retrieval
H.3.3; H.5.2; H.5.1
Multimodal deep-learning models power interactive video retrieval by ranking keyframes in response to textual queries. Despite these advances, users must still browse ranked candidates manually to locate a target. Keyframe arrangement within the search grid highly affects browsing effectiveness and user efficiency, yet remains underexplored. We report a study with 49 participants evaluating seven keyframe layouts for the Visual Known-Item Search task. Beyond efficiency and accuracy, we relate browsing phenomena, such as overlooks, to layout characteristics. Our results show that a video-grouped layout is the most efficient, while a four-column, rank-preserving grid achieves the highest accuracy. Sorted grids reveal potentials and trade-offs, enabling rapid scanning of uninteresting regions but down-ranking relevant targets to less prominent positions, delaying first arrival times and increasing overlooks. These findings motivate hybrid designs that preserve positions of top-ranked items while sorting or grouping the remainder, and offer guidance for searching in grids beyond video retrieval.
title Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections
topic Multimedia
Information Retrieval
H.3.3; H.5.2; H.5.1
url https://arxiv.org/abs/2510.04396