Saved in:
Bibliographic Details
Main Authors: Hwangbo, Junhyeong, Lee, Soohyun, Jeon, Hyeon, Jang, Kyochul, Cheong, Minsoo, Yu, Youngjae, Seo, Jinwook
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.10234
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Even LLMs that appear safe during evaluation can still produce harmful responses in deployment. Because stochastic sampling yields different responses to the same prompt, low-probability harmful outputs can still reach users at scale. Common human evaluation workflows generate many random samples per prompt and review them in static spreadsheets. The practice scales poorly, forcing evaluators to repeatedly reread near-duplicate prefixes. To address this, we present InFerActive, an interactive system that visualizes sampling results as a navigable tree of readable phrases, allowing evaluators to filter, explore, and extend the generation space on demand. InFerActive utilizes breadth-first sampling, a novel tree construction procedure that matches the harmful-response coverage of random sampling while requiring up to 5.0x fewer samples. Two controlled user studies (N = 12 each) demonstrate that InFerActive significantly improves evaluation efficiency and coverage over both spreadsheet and basic tree baselines.