STACT-Time: Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Adam, Irsyad, Zhang, Tengyue, Raman, Shrayes, Qiu, Zhuyu, Taraku, Brandon, Feng, Hexiang, Wang, Sile, Radhachandran, Ashwath, Athreya, Shreeram, Ivezic, Vedrana, Ping, Peipei, Arnold, Corey, Speier, William
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916808825503744
author Adam, Irsyad
Zhang, Tengyue
Raman, Shrayes
Qiu, Zhuyu
Taraku, Brandon
Feng, Hexiang
Wang, Sile
Radhachandran, Ashwath
Athreya, Shreeram
Ivezic, Vedrana
Ping, Peipei
Arnold, Corey
Speier, William
author_facet Adam, Irsyad
Zhang, Tengyue
Raman, Shrayes
Qiu, Zhuyu
Taraku, Brandon
Feng, Hexiang
Wang, Sile
Radhachandran, Ashwath
Athreya, Shreeram
Ivezic, Vedrana
Ping, Peipei
Arnold, Corey
Speier, William
contents Thyroid cancer is among the most common cancers in the United States. Thyroid nodules are frequently detected through ultrasound (US) imaging, and some require further evaluation via fine-needle aspiration (FNA) biopsy. Despite its effectiveness, FNA often leads to unnecessary biopsies of benign nodules, causing patient discomfort and anxiety. To address this, the American College of Radiology Thyroid Imaging Reporting and Data System (TI-RADS) has been developed to reduce benign biopsies. However, such systems are limited by interobserver variability. Recent deep learning approaches have sought to improve risk stratification, but they often fail to utilize the rich temporal and spatial context provided by US cine clips, which contain dynamic global information and surrounding structural changes across various views. In this work, we propose the Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification (STACT-Time) model, a novel representation learning framework that integrates imaging features from US cine clips with features from segmentation masks automatically generated by a pretrained model. By leveraging self-attention and cross-attention mechanisms, our model captures the rich temporal and spatial context of US cine clips while enhancing feature representation through segmentation-guided learning. Our model improves malignancy prediction compared to state-of-the-art models, achieving a cross-validation precision of 0.91 (plus or minus 0.02) and an F1 score of 0.89 (plus or minus 0.02). By reducing unnecessary biopsies of benign nodules while maintaining high sensitivity for malignancy detection, our model has the potential to enhance clinical decision-making and improve patient outcomes.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18172
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STACT-Time: Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification
Adam, Irsyad
Zhang, Tengyue
Raman, Shrayes
Qiu, Zhuyu
Taraku, Brandon
Feng, Hexiang
Wang, Sile
Radhachandran, Ashwath
Athreya, Shreeram
Ivezic, Vedrana
Ping, Peipei
Arnold, Corey
Speier, William
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Thyroid cancer is among the most common cancers in the United States. Thyroid nodules are frequently detected through ultrasound (US) imaging, and some require further evaluation via fine-needle aspiration (FNA) biopsy. Despite its effectiveness, FNA often leads to unnecessary biopsies of benign nodules, causing patient discomfort and anxiety. To address this, the American College of Radiology Thyroid Imaging Reporting and Data System (TI-RADS) has been developed to reduce benign biopsies. However, such systems are limited by interobserver variability. Recent deep learning approaches have sought to improve risk stratification, but they often fail to utilize the rich temporal and spatial context provided by US cine clips, which contain dynamic global information and surrounding structural changes across various views. In this work, we propose the Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification (STACT-Time) model, a novel representation learning framework that integrates imaging features from US cine clips with features from segmentation masks automatically generated by a pretrained model. By leveraging self-attention and cross-attention mechanisms, our model captures the rich temporal and spatial context of US cine clips while enhancing feature representation through segmentation-guided learning. Our model improves malignancy prediction compared to state-of-the-art models, achieving a cross-validation precision of 0.91 (plus or minus 0.02) and an F1 score of 0.89 (plus or minus 0.02). By reducing unnecessary biopsies of benign nodules while maintaining high sensitivity for malignancy detection, our model has the potential to enhance clinical decision-making and improve patient outcomes.
title STACT-Time: Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18172