UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Shaker, Abdelrahman, Maaz, Muhammad, Rasheed, Hanoona, Khan, Salman, Yang, Ming-Hsuan, Khan, Fahad Shahbaz |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
di: Maaz, Muhammad, et al.
Pubblicazione: (2023)
di: Maaz, Muhammad, et al.
Pubblicazione: (2023)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
di: Maaz, Muhammad, et al.
Pubblicazione: (2025)
di: Maaz, Muhammad, et al.
Pubblicazione: (2025)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
di: Maaz, Muhammad, et al.
Pubblicazione: (2024)
di: Maaz, Muhammad, et al.
Pubblicazione: (2024)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2025)
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
di: Kunhimon, Shahina, et al.
Pubblicazione: (2023)
di: Kunhimon, Shahina, et al.
Pubblicazione: (2023)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
GLaMM: Pixel Grounding Large Multimodal Model
di: Rasheed, Hanoona, et al.
Pubblicazione: (2023)
di: Rasheed, Hanoona, et al.
Pubblicazione: (2023)
PALO: A Polyglot Large Multimodal Model for 5B People
di: Maaz, Muhammad, et al.
Pubblicazione: (2024)
di: Maaz, Muhammad, et al.
Pubblicazione: (2024)
GroupMamba: Efficient Group-Based Visual State Space Model
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2024)
DB-SAM: Delving into High Quality Universal Medical Image Segmentation
di: Qin, Chao, et al.
Pubblicazione: (2024)
di: Qin, Chao, et al.
Pubblicazione: (2024)
Language Guided Domain Generalized Medical Image Segmentation
di: Kunhimon, Shahina, et al.
Pubblicazione: (2024)
di: Kunhimon, Shahina, et al.
Pubblicazione: (2024)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
di: Thawakar, Omkar, et al.
Pubblicazione: (2023)
di: Thawakar, Omkar, et al.
Pubblicazione: (2023)
WorldCache: Content-Aware Caching for Accelerated Video World Models
di: Nawaz, Umair, et al.
Pubblicazione: (2026)
di: Nawaz, Umair, et al.
Pubblicazione: (2026)
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
di: Boudjoghra, Mohamed El Amine, et al.
Pubblicazione: (2024)
di: Boudjoghra, Mohamed El Amine, et al.
Pubblicazione: (2024)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
di: Dudhane, Akshay, et al.
Pubblicazione: (2024)
di: Dudhane, Akshay, et al.
Pubblicazione: (2024)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
di: Ahmad, Ghazi Shazan, et al.
Pubblicazione: (2025)
di: Ahmad, Ghazi Shazan, et al.
Pubblicazione: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
di: Wasim, Syed Talal, et al.
Pubblicazione: (2023)
di: Wasim, Syed Talal, et al.
Pubblicazione: (2023)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
di: Gani, Hanan, et al.
Pubblicazione: (2024)
di: Gani, Hanan, et al.
Pubblicazione: (2024)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
di: Kumar, Komal, et al.
Pubblicazione: (2025)
di: Kumar, Komal, et al.
Pubblicazione: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
di: Dharmasiri, Amaya, et al.
Pubblicazione: (2024)
di: Dharmasiri, Amaya, et al.
Pubblicazione: (2024)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
di: Heakl, Ahmed, et al.
Pubblicazione: (2026)
di: Heakl, Ahmed, et al.
Pubblicazione: (2026)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
di: Ashraf, Tajamul, et al.
Pubblicazione: (2025)
di: Ashraf, Tajamul, et al.
Pubblicazione: (2025)
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
di: Hanif, Asif, et al.
Pubblicazione: (2024)
di: Hanif, Asif, et al.
Pubblicazione: (2024)
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
di: Ding, Bonan, et al.
Pubblicazione: (2026)
di: Ding, Bonan, et al.
Pubblicazione: (2026)
Diversity Has Always Been There in Your Visual Autoregressive Models
di: Wang, Tong, et al.
Pubblicazione: (2025)
di: Wang, Tong, et al.
Pubblicazione: (2025)
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
di: Sheikh, Tooba Tehreem, et al.
Pubblicazione: (2025)
di: Sheikh, Tooba Tehreem, et al.
Pubblicazione: (2025)
LawDIS: Language-Window-based Controllable Dichotomous Image Segmentation
di: Yan, Xinyu, et al.
Pubblicazione: (2025)
di: Yan, Xinyu, et al.
Pubblicazione: (2025)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
di: Thawakar, Omkar, et al.
Pubblicazione: (2025)
di: Thawakar, Omkar, et al.
Pubblicazione: (2025)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
di: Noman, Mubashir, et al.
Pubblicazione: (2024)
di: Noman, Mubashir, et al.
Pubblicazione: (2024)
On Evaluating Adversarial Robustness of Volumetric Medical Segmentation Models
di: Malik, Hashmat Shadab, et al.
Pubblicazione: (2024)
di: Malik, Hashmat Shadab, et al.
Pubblicazione: (2024)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
di: Chen, Shiming, et al.
Pubblicazione: (2025)
di: Chen, Shiming, et al.
Pubblicazione: (2025)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
di: Chen, Shiming, et al.
Pubblicazione: (2025)
di: Chen, Shiming, et al.
Pubblicazione: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
di: Chen, Shiming, et al.
Pubblicazione: (2024)
di: Chen, Shiming, et al.
Pubblicazione: (2024)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2026)
di: Shaker, Abdelrahman, et al.
Pubblicazione: (2026)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
di: Luo, Ziyang, et al.
Pubblicazione: (2025)
di: Luo, Ziyang, et al.
Pubblicazione: (2025)
MLRU++: Multiscale Lightweight Residual UNETR++ with Attention for Efficient 3D Medical Image Segmentation
di: Yadav, Nand Kumar, et al.
Pubblicazione: (2025)
di: Yadav, Nand Kumar, et al.
Pubblicazione: (2025)
Dual Hyperspectral Mamba for Efficient Spectral Compressive Imaging
di: Dong, Jiahua, et al.
Pubblicazione: (2024)
di: Dong, Jiahua, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025) -
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
di: Maaz, Muhammad, et al.
Pubblicazione: (2023) -
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
di: Maaz, Muhammad, et al.
Pubblicazione: (2025) -
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
di: Maaz, Muhammad, et al.
Pubblicazione: (2024) -
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
di: Rasheed, Hanoona, et al.
Pubblicazione: (2025)