Future Slot Prediction for Unsupervised Object Discovery in Surgical Video
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908439749328896 |
|---|---|
| author | Liao, Guiqiu Jogan, Matjaz Hussing, Marcel Zhang, Edward Eaton, Eric Hashimoto, Daniel A. |
| author_facet | Liao, Guiqiu Jogan, Matjaz Hussing, Marcel Zhang, Edward Eaton, Eric Hashimoto, Daniel A. |
| contents | Object-centric slot attention is an emerging paradigm for unsupervised learning of structured, interpretable object-centric representations (slots). This enables effective reasoning about objects and events at a low computational cost and is thus applicable to critical healthcare applications, such as real-time interpretation of surgical video. The heterogeneous scenes in real-world applications like surgery are, however, difficult to parse into a meaningful set of slots. Current approaches with an adaptive slot count perform well on images, but their performance on surgical videos is low. To address this challenge, we propose a dynamic temporal slot transformer (DTST) module that is trained both for temporal reasoning and for predicting the optimal future slot initialization. The model achieves state-of-the-art performance on multiple surgical databases, demonstrating that unsupervised object-centric methods can be applied to real-world data and become part of the common arsenal in healthcare applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_01882 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Future Slot Prediction for Unsupervised Object Discovery in Surgical Video Liao, Guiqiu Jogan, Matjaz Hussing, Marcel Zhang, Edward Eaton, Eric Hashimoto, Daniel A. Computer Vision and Pattern Recognition Object-centric slot attention is an emerging paradigm for unsupervised learning of structured, interpretable object-centric representations (slots). This enables effective reasoning about objects and events at a low computational cost and is thus applicable to critical healthcare applications, such as real-time interpretation of surgical video. The heterogeneous scenes in real-world applications like surgery are, however, difficult to parse into a meaningful set of slots. Current approaches with an adaptive slot count perform well on images, but their performance on surgical videos is low. To address this challenge, we propose a dynamic temporal slot transformer (DTST) module that is trained both for temporal reasoning and for predicting the optimal future slot initialization. The model achieves state-of-the-art performance on multiple surgical databases, demonstrating that unsupervised object-centric methods can be applied to real-world data and become part of the common arsenal in healthcare applications. |
| title | Future Slot Prediction for Unsupervised Object Discovery in Surgical Video |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2507.01882 |