Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Meng, Mingyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
Scalable Diffusion Models with State Space Backbone
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
RetFiner: A Vision-Language Refinement Scheme for Retinal Foundation Models
von: Fecso, Ronald, et al.
Veröffentlicht: (2025)
von: Fecso, Ronald, et al.
Veröffentlicht: (2025)
Vision KAN: Towards an Attention-Free Backbone for Vision with Kolmogorov-Arnold Networks
von: Yang, Zhuoqin, et al.
Veröffentlicht: (2026)
von: Yang, Zhuoqin, et al.
Veröffentlicht: (2026)
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
von: Ye, Shuchang, et al.
Veröffentlicht: (2025)
von: Ye, Shuchang, et al.
Veröffentlicht: (2025)
vGamba: Attentive State Space Bottleneck for efficient Long-range Dependencies in Visual Recognition
von: Haruna, Yunusa, et al.
Veröffentlicht: (2025)
von: Haruna, Yunusa, et al.
Veröffentlicht: (2025)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
von: Yin, Weijie, et al.
Veröffentlicht: (2025)
von: Yin, Weijie, et al.
Veröffentlicht: (2025)
Sparsity- and Hybridity-Inspired Visual Parameter-Efficient Fine-Tuning for Medical Diagnosis
von: Liu, Mingyuan, et al.
Veröffentlicht: (2024)
von: Liu, Mingyuan, et al.
Veröffentlicht: (2024)
MambaVision: A Hybrid Mamba-Transformer Vision Backbone
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2024)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2024)
Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis
von: Qu, Jingguo, et al.
Veröffentlicht: (2025)
von: Qu, Jingguo, et al.
Veröffentlicht: (2025)
Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation
von: Ye, Shuchang, et al.
Veröffentlicht: (2025)
von: Ye, Shuchang, et al.
Veröffentlicht: (2025)
Seedream 4.0: Toward Next-generation Multimodal Image Generation
von: Seedream, Team, et al.
Veröffentlicht: (2025)
von: Seedream, Team, et al.
Veröffentlicht: (2025)
The Finer the Better: Towards Granular-aware Open-set Domain Generalization
von: Wang, Yunyun, et al.
Veröffentlicht: (2025)
von: Wang, Yunyun, et al.
Veröffentlicht: (2025)
Parsing Objects at a Finer Granularity: A Survey
von: Zhao, Yifan, et al.
Veröffentlicht: (2022)
von: Zhao, Yifan, et al.
Veröffentlicht: (2022)
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
von: Ge, Chunjiang, et al.
Veröffentlicht: (2024)
von: Ge, Chunjiang, et al.
Veröffentlicht: (2024)
NextAds: Towards Next-generation Personalized Video Advertising
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
SegMamba: Long-range Sequential Modeling Mamba For 3D Medical Image Segmentation
von: Xing, Zhaohu, et al.
Veröffentlicht: (2024)
von: Xing, Zhaohu, et al.
Veröffentlicht: (2024)
ViR: Towards Efficient Vision Retention Backbones
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
Dynamic Traceback Learning for Medical Report Generation
von: Ye, Shuchang, et al.
Veröffentlicht: (2024)
von: Ye, Shuchang, et al.
Veröffentlicht: (2024)
There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks
von: Espinosa, Miguel, et al.
Veröffentlicht: (2024)
von: Espinosa, Miguel, et al.
Veröffentlicht: (2024)
Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation
von: Wei, Mingjie, et al.
Veröffentlicht: (2025)
von: Wei, Mingjie, et al.
Veröffentlicht: (2025)
HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone
von: Fu, Guanyiman, et al.
Veröffentlicht: (2026)
von: Fu, Guanyiman, et al.
Veröffentlicht: (2026)
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
von: Park, Woohyeon, et al.
Veröffentlicht: (2026)
von: Park, Woohyeon, et al.
Veröffentlicht: (2026)
Finer Disentanglement of Aleatoric Uncertainty Can Accelerate Chemical Histopathology Imaging
von: Oh, Ji-Hun, et al.
Veröffentlicht: (2025)
von: Oh, Ji-Hun, et al.
Veröffentlicht: (2025)
VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
von: Bose, Sarosij, et al.
Veröffentlicht: (2025)
von: Bose, Sarosij, et al.
Veröffentlicht: (2025)
Revisiting the Integration of Convolution and Attention for Vision Backbone
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
von: Safdar, Aon, et al.
Veröffentlicht: (2025)
von: Safdar, Aon, et al.
Veröffentlicht: (2025)
ImplantMamba: Long-range Sequential Modeling Mamba For Dental Implant Position Prediction
von: Yang, Xinquan, et al.
Veröffentlicht: (2026)
von: Yang, Xinquan, et al.
Veröffentlicht: (2026)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
Language-guided Medical Image Segmentation with Target-informed Multi-level Contrastive Alignments
von: Li, Mingjian, et al.
Veröffentlicht: (2024)
von: Li, Mingjian, et al.
Veröffentlicht: (2024)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
SemRaFiner: Panoptic Segmentation in Sparse and Noisy Radar Point Clouds
von: Zeller, Matthias, et al.
Veröffentlicht: (2025)
von: Zeller, Matthias, et al.
Veröffentlicht: (2025)
MP-PolarMask: A Faster and Finer Instance Segmentation for Concave Images
von: Wang, Ke-Lei, et al.
Veröffentlicht: (2024)
von: Wang, Ke-Lei, et al.
Veröffentlicht: (2024)
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
von: An, Xiang, et al.
Veröffentlicht: (2026)
von: An, Xiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025) -
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024) -
Scalable Diffusion Models with State Space Backbone
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024) -
RetFiner: A Vision-Language Refinement Scheme for Retinal Foundation Models
von: Fecso, Ronald, et al.
Veröffentlicht: (2025) -
Vision KAN: Towards an Attention-Free Backbone for Vision with Kolmogorov-Arnold Networks
von: Yang, Zhuoqin, et al.
Veröffentlicht: (2026)