ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xucheng, Zhang, Xiaoman, Kim, Sung Eun, Pal, Ankit, Rajpurkar, Pranav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding
by: Pal, Ankit, et al.
Published: (2025)
by: Pal, Ankit, et al.
Published: (2025)
ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
by: Kenia, Roshan, et al.
Published: (2025)
by: Kenia, Roshan, et al.
Published: (2025)
3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
by: Sambara, Sraavya, et al.
Published: (2025)
by: Sambara, Sraavya, et al.
Published: (2025)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
by: Zhou, Zhiyu, et al.
Published: (2026)
by: Zhou, Zhiyu, et al.
Published: (2026)
Uncovering Knowledge Gaps in Radiology Report Generation Models through Knowledge Graphs
by: Zhang, Xiaoman, et al.
Published: (2024)
by: Zhang, Xiaoman, et al.
Published: (2024)
FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models
by: Heiman, Alice, et al.
Published: (2024)
by: Heiman, Alice, et al.
Published: (2024)
ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors
by: Hardy, Romain, et al.
Published: (2025)
by: Hardy, Romain, et al.
Published: (2025)
Evaluating Contextual Intelligence in Recyclability: A Comprehensive Study of Image-Based Reasoning Systems
by: Park, Eliot, et al.
Published: (2025)
by: Park, Eliot, et al.
Published: (2025)
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
by: Zhang, Xiaoman, et al.
Published: (2024)
by: Zhang, Xiaoman, et al.
Published: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
by: Kou, Qian, et al.
Published: (2026)
by: Kou, Qian, et al.
Published: (2026)
ReXInTheWild: A Unified Benchmark for Medical Photograph Understanding
by: Banerjee, Oishi, et al.
Published: (2026)
by: Banerjee, Oishi, et al.
Published: (2026)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
by: Kim, Kangsan, et al.
Published: (2026)
by: Kim, Kangsan, et al.
Published: (2026)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
by: Baharoon, Mohammed, et al.
Published: (2026)
by: Baharoon, Mohammed, et al.
Published: (2026)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026)
by: Durante, Zane, et al.
Published: (2026)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
by: Yu, Suhao, et al.
Published: (2025)
by: Yu, Suhao, et al.
Published: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
by: Li, Chenglin, et al.
Published: (2024)
by: Li, Chenglin, et al.
Published: (2024)
TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
by: Kim, Yoonsik, et al.
Published: (2024)
by: Kim, Yoonsik, et al.
Published: (2024)
DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
by: Yu, Xiaoxuan, et al.
Published: (2024)
by: Yu, Xiaoxuan, et al.
Published: (2024)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
by: Yang, Zhuoyi, et al.
Published: (2026)
by: Yang, Zhuoyi, et al.
Published: (2026)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
KNVQA: A Benchmark for evaluation knowledge-based VQA
by: Cheng, Sirui, et al.
Published: (2023)
by: Cheng, Sirui, et al.
Published: (2023)
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
by: Kim, Kangsan, et al.
Published: (2024)
by: Kim, Kangsan, et al.
Published: (2024)
a2z-1 for Multi-Disease Detection in Abdomen-Pelvis CT: External Validation and Performance Analysis Across 21 Conditions
by: Rajpurkar, Pranav, et al.
Published: (2024)
by: Rajpurkar, Pranav, et al.
Published: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
by: Ma, Jianzhe, et al.
Published: (2026)
by: Ma, Jianzhe, et al.
Published: (2026)
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
by: Gan, Rui, et al.
Published: (2026)
by: Gan, Rui, et al.
Published: (2026)
NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models
by: Park, Sung-Yeon, et al.
Published: (2025)
by: Park, Sung-Yeon, et al.
Published: (2025)
R^3-VQA: "Read the Room" by Video Social Reasoning
by: Niu, Lixing, et al.
Published: (2025)
by: Niu, Lixing, et al.
Published: (2025)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
YTCommentQA: Video Question Answerability in Instructional Videos
by: Yang, Saelyne, et al.
Published: (2024)
by: Yang, Saelyne, et al.
Published: (2024)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
by: Song, Yeon-Ji, et al.
Published: (2024)
by: Song, Yeon-Ji, et al.
Published: (2024)
ReWind: Understanding Long Videos with Instructed Learnable Memory
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
by: Ge, Qihang, et al.
Published: (2024)
by: Ge, Qihang, et al.
Published: (2024)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2025)
by: Seth, Ashish, et al.
Published: (2025)
LVBench: An Extreme Long Video Understanding Benchmark
by: Wang, Weihan, et al.
Published: (2024)
by: Wang, Weihan, et al.
Published: (2024)
Similar Items
-
ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding
by: Pal, Ankit, et al.
Published: (2025) -
ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
by: Kenia, Roshan, et al.
Published: (2025) -
3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
by: Sambara, Sraavya, et al.
Published: (2025) -
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
by: Zhou, Zhiyu, et al.
Published: (2026) -
Uncovering Knowledge Gaps in Radiology Report Generation Models through Knowledge Graphs
by: Zhang, Xiaoman, et al.
Published: (2024)