Representation geometry shapes task performance in vision-language modeling for CT enterography
Fuente:
arXiv
Saved in:
| Main Authors: | Minoccheri, Cristian, Wittrup, Emily, Najarian, Kayvan, Stidham, Ryan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoRA-based methods on Unet for transfer learning in Subarachnoid Hematoma Segmentation
by: Minoccheri, Cristian, et al.
Published: (2025)
by: Minoccheri, Cristian, et al.
Published: (2025)
A Causal Machine Learning Framework for Treatment Personalization in Clinical Trials: Application to Ulcerative Colitis
by: Minoccheri, Cristian, et al.
Published: (2026)
by: Minoccheri, Cristian, et al.
Published: (2026)
Supervised Coupled Matrix-Tensor Factorization (SCMTF) for Computational Phenotyping of Patient Reported Outcomes in Ulcerative Colitis
by: Minoccheri, Cristian, et al.
Published: (2025)
by: Minoccheri, Cristian, et al.
Published: (2025)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026)
by: Pan, Baiyu, et al.
Published: (2026)
Hallucination-aware intermediate representation edit in large vision-language models
by: Suo, Wei, et al.
Published: (2026)
by: Suo, Wei, et al.
Published: (2026)
Building and better understanding vision-language models: insights and future directions
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Generalizing vision-language models to novel domains: A comprehensive survey
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
by: Nan, Yang, et al.
Published: (2024)
by: Nan, Yang, et al.
Published: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
by: Guo, Xiaotong, et al.
Published: (2025)
by: Guo, Xiaotong, et al.
Published: (2025)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Improving vision-language alignment with graph spiking hybrid Networks
by: Zhang, Siyu, et al.
Published: (2025)
by: Zhang, Siyu, et al.
Published: (2025)
Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models
by: Aman, Tabinda, et al.
Published: (2025)
by: Aman, Tabinda, et al.
Published: (2025)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
by: Dong, Owen, et al.
Published: (2026)
by: Dong, Owen, et al.
Published: (2026)
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images
by: Meseguer, Pablo, et al.
Published: (2024)
by: Meseguer, Pablo, et al.
Published: (2024)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
Evaluating Large Vision-language Models for Surgical Tool Detection
by: Poudel, Nakul, et al.
Published: (2026)
by: Poudel, Nakul, et al.
Published: (2026)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
by: Walimbe, Soham, et al.
Published: (2025)
by: Walimbe, Soham, et al.
Published: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
Implicit Neural Representations for Robust Joint Sparse-View CT Reconstruction
by: Shi, Jiayang, et al.
Published: (2024)
by: Shi, Jiayang, et al.
Published: (2024)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
Teaching large language models to reason like expert diagnosticians
by: Buckley, Thomas A., et al.
Published: (2025)
by: Buckley, Thomas A., et al.
Published: (2025)
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
by: Kim, Daeyoung
Published: (2026)
by: Kim, Daeyoung
Published: (2026)
Vision-language models lag human performance on physical dynamics and intent reasoning
by: Gu, Tianjun, et al.
Published: (2026)
by: Gu, Tianjun, et al.
Published: (2026)
Towards aligned body representations in vision models
by: Gizdov, Andrey, et al.
Published: (2025)
by: Gizdov, Andrey, et al.
Published: (2025)
Advancing vision-language models in front-end development via data synthesis
by: Ge, Tong, et al.
Published: (2025)
by: Ge, Tong, et al.
Published: (2025)
PathAlign: A vision-language model for whole slide images in histopathology
by: Ahmed, Faruk, et al.
Published: (2024)
by: Ahmed, Faruk, et al.
Published: (2024)
Vision language models are unreliable at trivial spatial cognition
by: Khemlani, Sangeet, et al.
Published: (2025)
by: Khemlani, Sangeet, et al.
Published: (2025)
Prompting with the human-touch: evaluating model-sensitivity of foundation models for musculoskeletal CT segmentation
by: Magg, Caroline, et al.
Published: (2026)
by: Magg, Caroline, et al.
Published: (2026)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Vision language models have difficulty recognizing virtual objects
by: Tran, Tyler, et al.
Published: (2025)
by: Tran, Tyler, et al.
Published: (2025)
See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning
by: Zheng, Chengxin, et al.
Published: (2024)
by: Zheng, Chengxin, et al.
Published: (2024)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
Evaluating point-light biological motion in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
Multi LoRA Meets Vision: Merging multiple adapters to create a multi task model
by: Kesim, Ege, et al.
Published: (2024)
by: Kesim, Ege, et al.
Published: (2024)
ViSTa Dataset: Do vision-language models understand sequential tasks?
by: Wybitul, Evžen, et al.
Published: (2024)
by: Wybitul, Evžen, et al.
Published: (2024)
Similar Items
-
LoRA-based methods on Unet for transfer learning in Subarachnoid Hematoma Segmentation
by: Minoccheri, Cristian, et al.
Published: (2025) -
A Causal Machine Learning Framework for Treatment Personalization in Clinical Trials: Application to Ulcerative Colitis
by: Minoccheri, Cristian, et al.
Published: (2026) -
Supervised Coupled Matrix-Tensor Factorization (SCMTF) for Computational Phenotyping of Patient Reported Outcomes in Ulcerative Colitis
by: Minoccheri, Cristian, et al.
Published: (2025) -
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025) -
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)