ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xucheng, Zhang, Xiaoman, Kim, Sung Eun, Pal, Ankit, Rajpurkar, Pranav
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915941622743040
author Wang, Xucheng
Zhang, Xiaoman
Kim, Sung Eun
Pal, Ankit
Rajpurkar, Pranav
author_facet Wang, Xucheng
Zhang, Xiaoman
Kim, Sung Eun
Pal, Ankit
Rajpurkar, Pranav
contents Ultrasound acquisition requires skilled probe manipulation and real-time adjustments. Vision-language models (VLMs) could enable autonomous ultrasound systems, but existing benchmarks evaluate only static images, not dynamic procedural understanding. We introduce ReXSonoVQA, a video QA benchmark with 514 video clips and 514 questions (249 MCQ, 265 free-response) targeting three competencies: Action-Goal Reasoning, Artifact Resolution & Optimization, and Procedure Context & Planning. Zero-shot evaluation of Gemini 3 Pro, Qwen3.5-397B, LLaVA-Video-72B, and Seed 2.0 Pro shows VLMs can extract some procedural information, but troubleshooting questions remain challenging with minimal gains over text-only baselines, exposing limitations in causal reasoning. ReXSonoVQA enables developing perception systems for ultrasound training, guidance, and robotic automation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10916
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
Wang, Xucheng
Zhang, Xiaoman
Kim, Sung Eun
Pal, Ankit
Rajpurkar, Pranav
Computer Vision and Pattern Recognition
Artificial Intelligence
Ultrasound acquisition requires skilled probe manipulation and real-time adjustments. Vision-language models (VLMs) could enable autonomous ultrasound systems, but existing benchmarks evaluate only static images, not dynamic procedural understanding. We introduce ReXSonoVQA, a video QA benchmark with 514 video clips and 514 questions (249 MCQ, 265 free-response) targeting three competencies: Action-Goal Reasoning, Artifact Resolution & Optimization, and Procedure Context & Planning. Zero-shot evaluation of Gemini 3 Pro, Qwen3.5-397B, LLaVA-Video-72B, and Seed 2.0 Pro shows VLMs can extract some procedural information, but troubleshooting questions remain challenging with minimal gains over text-only baselines, exposing limitations in causal reasoning. ReXSonoVQA enables developing perception systems for ultrasound training, guidance, and robotic automation.
title ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.10916