Lunguage: A Benchmark for Structured and Sequential Chest X-ray Interpretation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Moon, Jong Hak, Choi, Geon, Rabaey, Paloma, Kim, Min Gwan, Lee, Jung-Oh, Hong, Hyuk Gi, Doe, Eun Woo, Yoon, Hangyul, Kim, Jiyoun, Sharma, Harshita, Castro, Daniel C., Alvarez-Valle, Javier, Choi, Edward
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908999186644992
author Moon, Jong Hak
Choi, Geon
Rabaey, Paloma
Kim, Min Gwan
Lee, Jung-Oh
Hong, Hyuk Gi
Doe, Eun Woo
Yoon, Hangyul
Kim, Jiyoun
Sharma, Harshita
Castro, Daniel C.
Alvarez-Valle, Javier
Choi, Edward
author_facet Moon, Jong Hak
Choi, Geon
Rabaey, Paloma
Kim, Min Gwan
Lee, Jung-Oh
Hong, Hyuk Gi
Doe, Eun Woo
Yoon, Hangyul
Kim, Jiyoun
Sharma, Harshita
Castro, Daniel C.
Alvarez-Valle, Javier
Choi, Edward
contents Radiology reports convey detailed clinical observations and capture diagnostic reasoning that evolves over time. However, existing evaluation methods are limited to single-report settings and rely on coarse metrics that fail to capture fine-grained clinical semantics and temporal dependencies. We introduce LUNGUAGE, a benchmark dataset for structured radiology report generation that supports both single-report evaluation and longitudinal patient-level assessment across multiple studies. It contains 1,473 annotated chest X-ray reports, each reviewed by experts, and 186 of them contain longitudinal annotations to capture disease progression and inter-study intervals, also reviewed by experts. Using this benchmark, we develop a two-stage structuring framework that transforms generated reports into fine-grained, schema-aligned structured reports, enabling longitudinal interpretation. We also propose LUNGUAGESCORE, an interpretable metric that compares structured outputs at the entity, relation, and attribute level while modeling temporal consistency across patient timelines. These contributions establish the first benchmark dataset, structuring framework, and evaluation metric for sequential radiology reporting, with empirical results demonstrating that LUNGUAGESCORE effectively supports structured report evaluation. The code is available at: https://github.com/SuperSupermoon/Lunguage
format Preprint
id arxiv_https___arxiv_org_abs_2505_21190
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lunguage: A Benchmark for Structured and Sequential Chest X-ray Interpretation
Moon, Jong Hak
Choi, Geon
Rabaey, Paloma
Kim, Min Gwan
Lee, Jung-Oh
Hong, Hyuk Gi
Doe, Eun Woo
Yoon, Hangyul
Kim, Jiyoun
Sharma, Harshita
Castro, Daniel C.
Alvarez-Valle, Javier
Choi, Edward
Computation and Language
Artificial Intelligence
Radiology reports convey detailed clinical observations and capture diagnostic reasoning that evolves over time. However, existing evaluation methods are limited to single-report settings and rely on coarse metrics that fail to capture fine-grained clinical semantics and temporal dependencies. We introduce LUNGUAGE, a benchmark dataset for structured radiology report generation that supports both single-report evaluation and longitudinal patient-level assessment across multiple studies. It contains 1,473 annotated chest X-ray reports, each reviewed by experts, and 186 of them contain longitudinal annotations to capture disease progression and inter-study intervals, also reviewed by experts. Using this benchmark, we develop a two-stage structuring framework that transforms generated reports into fine-grained, schema-aligned structured reports, enabling longitudinal interpretation. We also propose LUNGUAGESCORE, an interpretable metric that compares structured outputs at the entity, relation, and attribute level while modeling temporal consistency across patient timelines. These contributions establish the first benchmark dataset, structuring framework, and evaluation metric for sequential radiology reporting, with empirical results demonstrating that LUNGUAGESCORE effectively supports structured report evaluation. The code is available at: https://github.com/SuperSupermoon/Lunguage
title Lunguage: A Benchmark for Structured and Sequential Chest X-ray Interpretation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.21190