From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yuan, Kun, Sun, Min Woo, Chen, Zhen, Lozano, Alejandro, He, Xiangteng, Li, Shi, Navab, Nassir, Sun, Xiaoxiao, Padoy, Nicolas, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
par: Yuan, Kun, et autres
Publié: (2024)
par: Yuan, Kun, et autres
Publié: (2024)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
par: Yuan, Kun, et autres
Publié: (2024)
par: Yuan, Kun, et autres
Publié: (2024)
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
par: Stilz, Florian, et autres
Publié: (2026)
par: Stilz, Florian, et autres
Publié: (2026)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
par: Sun, Min Woo, et autres
Publié: (2025)
par: Sun, Min Woo, et autres
Publié: (2025)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
par: Lozano, Alejandro, et autres
Publié: (2025)
par: Lozano, Alejandro, et autres
Publié: (2025)
Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis
par: Chen, Tingxuan, et autres
Publié: (2025)
par: Chen, Tingxuan, et autres
Publié: (2025)
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
par: Lozano, Alejandro, et autres
Publié: (2025)
par: Lozano, Alejandro, et autres
Publié: (2025)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
par: Burgess, James, et autres
Publié: (2026)
par: Burgess, James, et autres
Publié: (2026)
Advancing Surgical VQA with Scene Graph Knowledge
par: Yuan, Kun, et autres
Publié: (2023)
par: Yuan, Kun, et autres
Publié: (2023)
Can Large Language Models Match the Conclusions of Systematic Reviews?
par: Polzak, Christopher, et autres
Publié: (2025)
par: Polzak, Christopher, et autres
Publié: (2025)
Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment
par: Shen, Yaling, et autres
Publié: (2025)
par: Shen, Yaling, et autres
Publié: (2025)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
par: Zhou, Yue, et autres
Publié: (2026)
par: Zhou, Yue, et autres
Publié: (2026)
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
par: Sun, Xiaoxiao, et autres
Publié: (2026)
par: Sun, Xiaoxiao, et autres
Publié: (2026)
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
par: Yuan, Kun, et autres
Publié: (2023)
par: Yuan, Kun, et autres
Publié: (2023)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
par: Hou, Wenjin, et autres
Publié: (2026)
par: Hou, Wenjin, et autres
Publié: (2026)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
par: Aklilu, Josiah, et autres
Publié: (2024)
par: Aklilu, Josiah, et autres
Publié: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
par: Endo, Mark, et autres
Publié: (2024)
par: Endo, Mark, et autres
Publié: (2024)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
par: Sui, Elaine, et autres
Publié: (2024)
par: Sui, Elaine, et autres
Publié: (2024)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
par: He, Xiangteng, et autres
Publié: (2025)
par: He, Xiangteng, et autres
Publié: (2025)
Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion
par: Wei, Meng, et autres
Publié: (2026)
par: Wei, Meng, et autres
Publié: (2026)
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
par: Huang, Yiming, et autres
Publié: (2025)
par: Huang, Yiming, et autres
Publié: (2025)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
par: Chen, Liangyu, et autres
Publié: (2025)
par: Chen, Liangyu, et autres
Publié: (2025)
Dyadic Partnership(DP): A Missing Link Towards Full Autonomy in Medical Robotics
par: Navab, Nassir, et autres
Publié: (2026)
par: Navab, Nassir, et autres
Publié: (2026)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
par: Lozano, Alejandro, et autres
Publié: (2024)
par: Lozano, Alejandro, et autres
Publié: (2024)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
par: Endo, Mark, et autres
Publié: (2025)
par: Endo, Mark, et autres
Publié: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
par: Gu, Jeffrey, et autres
Publié: (2025)
par: Gu, Jeffrey, et autres
Publié: (2025)
NegVQA: Can Vision Language Models Understand Negation?
par: Zhang, Yuhui, et autres
Publié: (2025)
par: Zhang, Yuhui, et autres
Publié: (2025)
ORacle: Large Vision-Language Models for Knowledge-Guided Holistic OR Domain Modeling
par: Özsoy, Ege, et autres
Publié: (2024)
par: Özsoy, Ege, et autres
Publié: (2024)
Uncertainty-Aware Distribution-to-Distribution Flow Matching for Scientific Imaging
par: Wu, Dongxia, et autres
Publié: (2026)
par: Wu, Dongxia, et autres
Publié: (2026)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
par: Heo, Jaewoo, et autres
Publié: (2024)
par: Heo, Jaewoo, et autres
Publié: (2024)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
par: Hu, Ming, et autres
Publié: (2024)
par: Hu, Ming, et autres
Publié: (2024)
Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
par: Yuan, Kun, et autres
Publié: (2025)
par: Yuan, Kun, et autres
Publié: (2025)
Revisiting Active Learning in the Era of Vision Foundation Models
par: Gupte, Sanket Rajan, et autres
Publié: (2024)
par: Gupte, Sanket Rajan, et autres
Publié: (2024)
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
par: Biagini, Diego, et autres
Publié: (2025)
par: Biagini, Diego, et autres
Publié: (2025)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
par: Holm, Felix, et autres
Publié: (2025)
par: Holm, Felix, et autres
Publié: (2025)
SURGIVID: Annotation-Efficient Surgical Video Object Discovery
par: Köksal, Çağhan, et autres
Publié: (2024)
par: Köksal, Çağhan, et autres
Publié: (2024)
Improving Robustness to Out-of-Distribution States in Imitation Learning via Deep Koopman-Boosted Diffusion Policy
par: Huang, Dianye, et autres
Publié: (2025)
par: Huang, Dianye, et autres
Publié: (2025)
Speckle2Self: Self-Supervised Ultrasound Speckle Reduction Without Clean Data
par: Li, Xuesong, et autres
Publié: (2025)
par: Li, Xuesong, et autres
Publié: (2025)
Neural Semantic Map-Learning for Autonomous Vehicles
par: Herb, Markus, et autres
Publié: (2024)
par: Herb, Markus, et autres
Publié: (2024)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
par: Sharma, Saurav, et autres
Publié: (2025)
par: Sharma, Saurav, et autres
Publié: (2025)
Documents similaires
-
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
par: Yuan, Kun, et autres
Publié: (2024) -
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
par: Yuan, Kun, et autres
Publié: (2024) -
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
par: Stilz, Florian, et autres
Publié: (2026) -
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
par: Sun, Min Woo, et autres
Publié: (2025) -
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
par: Lozano, Alejandro, et autres
Publié: (2025)