Learning complete and explainable visual representations from itemized text supervision
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Yiwei, Zhao, Chenhui, Banerjee, Soumyanil, Liu, Shixuan, Rao, Akshay, Kondepudi, Akhil, Lee, Honglak, Hollon, Todd C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A self-supervised framework for learning whole slide representations
by: Hou, Xinhai, et al.
Published: (2024)
by: Hou, Xinhai, et al.
Published: (2024)
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
by: Zhao, Chenhui, et al.
Published: (2025)
by: Zhao, Chenhui, et al.
Published: (2025)
Super-resolution of biomedical volumes with 2D supervision
by: Jiang, Cheng, et al.
Published: (2024)
by: Jiang, Cheng, et al.
Published: (2024)
Learning neuroimaging models from health system-scale data
by: Lyu, Yiwei, et al.
Published: (2025)
by: Lyu, Yiwei, et al.
Published: (2025)
Health system learning achieves generalist neuroimaging models
by: Kondepudi, Akhil, et al.
Published: (2025)
by: Kondepudi, Akhil, et al.
Published: (2025)
Step-Calibrated Diffusion for Biomedical Optical Image Restoration
by: Lyu, Yiwei, et al.
Published: (2024)
by: Lyu, Yiwei, et al.
Published: (2024)
Extending SEEDS to a Supervoxel Algorithm for Medical Image Analysis
by: Zhao, Chenhui, et al.
Published: (2025)
by: Zhao, Chenhui, et al.
Published: (2025)
Volumetric medical image segmentation through dual self-distillation in U-shaped networks
by: Banerjee, Soumyanil, et al.
Published: (2023)
by: Banerjee, Soumyanil, et al.
Published: (2023)
Conditional diffusion model with spatial attention and latent embedding for medical image segmentation
by: Hejrati, Behzad, et al.
Published: (2025)
by: Hejrati, Behzad, et al.
Published: (2025)
Do text-free diffusion models learn discriminative visual representations?
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
Investigating self-supervised representations for audio-visual deepfake detection
by: Boldisor, Dragos-Alexandru, et al.
Published: (2025)
by: Boldisor, Dragos-Alexandru, et al.
Published: (2025)
View Selection for 3D Captioning via Diffusion Ranking
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
by: Hou, Xinhai, et al.
Published: (2025)
by: Hou, Xinhai, et al.
Published: (2025)
Weakly-supervised Camera Localization by Ground-to-satellite Image Registration
by: Shi, Yujiao, et al.
Published: (2024)
by: Shi, Yujiao, et al.
Published: (2024)
Exploring scalable medical image encoders beyond text supervision
by: Pérez-García, Fernando, et al.
Published: (2024)
by: Pérez-García, Fernando, et al.
Published: (2024)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
Significantly improving zero-shot X-ray pathology classification via fine-tuning pre-trained image-text encoders
by: Jang, Jongseong, et al.
Published: (2022)
by: Jang, Jongseong, et al.
Published: (2022)
Towards a text-based quantitative and explainable histopathology image analysis
by: Nguyen, Anh Tien, et al.
Published: (2024)
by: Nguyen, Anh Tien, et al.
Published: (2024)
Self-supervised structured object representation learning
by: Hadjerci, Oussama, et al.
Published: (2025)
by: Hadjerci, Oussama, et al.
Published: (2025)
Intelligent Histology for Tumor Neurosurgery
by: Hou, Xinhai, et al.
Published: (2025)
by: Hou, Xinhai, et al.
Published: (2025)
A self-supervised text-vision framework for automated brain abnormality detection
by: Wood, David A., et al.
Published: (2024)
by: Wood, David A., et al.
Published: (2024)
A Semi-supervised Physics-Aware Triple-Stream Underwater Image Enhancement Network
by: Xu, Shixuan, et al.
Published: (2023)
by: Xu, Shixuan, et al.
Published: (2023)
MIMRS: A Survey on Masked Image Modeling in Remote Sensing
by: Choudhury, Shabnam, et al.
Published: (2025)
by: Choudhury, Shabnam, et al.
Published: (2025)
AACP: Aesthetics assessment of children's paintings based on self-supervised learning
by: Jiang, Shiqi, et al.
Published: (2024)
by: Jiang, Shiqi, et al.
Published: (2024)
Universal dimensions of visual representation
by: Chen, Zirui, et al.
Published: (2024)
by: Chen, Zirui, et al.
Published: (2024)
Part-aware Prompted Segment Anything Model for Adaptive Segmentation
by: Zhao, Chenhui, et al.
Published: (2024)
by: Zhao, Chenhui, et al.
Published: (2024)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Affine transformation estimation improves visual self-supervised learning
by: Torpey, David, et al.
Published: (2024)
by: Torpey, David, et al.
Published: (2024)
AnthroTAP: Learning Point Tracking with Real-World Motion
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
Characterizing the visual representation of objects from the child's view
by: Yang, Jane, et al.
Published: (2026)
by: Yang, Jane, et al.
Published: (2026)
A transition towards virtual representations of visual scenes
by: Pereira, Américo, et al.
Published: (2024)
by: Pereira, Américo, et al.
Published: (2024)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Self-supervised visual learning from interactions with objects
by: Aubret, Arthur, et al.
Published: (2024)
by: Aubret, Arthur, et al.
Published: (2024)
Slimmable Networks for Contrastive Self-supervised Learning
by: Zhao, Shuai, et al.
Published: (2022)
by: Zhao, Shuai, et al.
Published: (2022)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Weakly supervised training of universal visual concepts for multi-domain semantic segmentation
by: Bevandić, Petra, et al.
Published: (2022)
by: Bevandić, Petra, et al.
Published: (2022)
Self-supervised visual learning in the low-data regime: a comparative evaluation
by: Konstantakos, Sotirios, et al.
Published: (2024)
by: Konstantakos, Sotirios, et al.
Published: (2024)
Hyperspectral and multispectral image fusion with arbitrary resolution through self-supervised representations
by: Wang, Ting, et al.
Published: (2024)
by: Wang, Ting, et al.
Published: (2024)
Audio-visual training for improved grounding in video-text LLMs
by: Sagare, Shivprasad, et al.
Published: (2024)
by: Sagare, Shivprasad, et al.
Published: (2024)
A supervised discriminant data representation: application to pattern classification
by: Dornaika, Fadi, et al.
Published: (2025)
by: Dornaika, Fadi, et al.
Published: (2025)
Similar Items
-
A self-supervised framework for learning whole slide representations
by: Hou, Xinhai, et al.
Published: (2024) -
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
by: Zhao, Chenhui, et al.
Published: (2025) -
Super-resolution of biomedical volumes with 2D supervision
by: Jiang, Cheng, et al.
Published: (2024) -
Learning neuroimaging models from health system-scale data
by: Lyu, Yiwei, et al.
Published: (2025) -
Health system learning achieves generalist neuroimaging models
by: Kondepudi, Akhil, et al.
Published: (2025)