How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
Fuente:
arXiv
Saved in:
| Main Authors: | Thapa, Rahul, Li, Andrew, Wu, Qingyang, He, Bryan, Sahashi, Yuki, Binder, Christina, Zhang, Angela, Athiwaratkun, Ben, Song, Shuaiwen Leon, Ouyang, David, Zou, James |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
by: Thapa, Rahul, et al.
Published: (2024)
by: Thapa, Rahul, et al.
Published: (2024)
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
by: Dokmeci, Berkan, et al.
Published: (2025)
by: Dokmeci, Berkan, et al.
Published: (2025)
EchoNet-Quality: Denoising Echocardiograms via Deep Generative Modeling of Ultrasound Noise
by: Choi, David, et al.
Published: (2025)
by: Choi, David, et al.
Published: (2025)
Automated Interpretable 2D Video Extraction from 3D Echocardiography
by: Vukadinovic, Milos, et al.
Published: (2025)
by: Vukadinovic, Milos, et al.
Published: (2025)
Disentangling Reasoning and Knowledge in Medical Large Language Models
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning
by: Li, Andrew, et al.
Published: (2025)
by: Li, Andrew, et al.
Published: (2025)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Effects of Periodontal Therapy on Blood Lipid Levels in 10 Dogs With Periodontitis and Hyperlipidemia
by: Yu Sahashi, et al.
Published: (2025)
by: Yu Sahashi, et al.
Published: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Introspective Diffusion Language Models
by: Yu, Yifan, et al.
Published: (2026)
by: Yu, Yifan, et al.
Published: (2026)
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient
by: Zhou, Zhongzhu, et al.
Published: (2025)
by: Zhou, Zhongzhu, et al.
Published: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
by: Oncescu, Costin-Andrei, et al.
Published: (2025)
by: Oncescu, Costin-Andrei, et al.
Published: (2025)
Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
by: Zhang, Muru, et al.
Published: (2025)
by: Zhang, Muru, et al.
Published: (2025)
How Rural Information Systems Initiatives for Institutional Change Can Aggravate Inequalities: An Affordance‐Based Institutional Logics Perspective
by: Pragyan Thapa, et al.
Published: (2025)
by: Pragyan Thapa, et al.
Published: (2025)
Mixture-of-Agents Enhances Large Language Model Capabilities
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
by: Jia, Jinda, et al.
Published: (2026)
by: Jia, Jinda, et al.
Published: (2026)
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
Toward Energy‐Efficient Machine Vision: Advances in Optoelectronic Memristors
by: Shuaiwen Pan, et al.
Published: (2025)
by: Shuaiwen Pan, et al.
Published: (2025)
When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
by: Wang, Junxiong, et al.
Published: (2026)
by: Wang, Junxiong, et al.
Published: (2026)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025)
by: Ramachandran, Rahul, et al.
Published: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals
by: Thapa, Rahul, et al.
Published: (2024)
by: Thapa, Rahul, et al.
Published: (2024)
How Well Can AI Build SD Models?
by: Schoenberg, William, et al.
Published: (2025)
by: Schoenberg, William, et al.
Published: (2025)
Origin of Reduced Coercive Field in ScAlN: Synergy of Structural Softening and Dynamic Atomic Correlations
by: Sahashi, Ryotaro, et al.
Published: (2026)
by: Sahashi, Ryotaro, et al.
Published: (2026)
Decoupling structural and bonding effects on ferroelectric switching in ScAlN via molecular dynamics under an applied electric field
by: Sahashi, Ryotaro, et al.
Published: (2026)
by: Sahashi, Ryotaro, et al.
Published: (2026)
Can Education for Sustainable Development Support Climate Change Adaptation Effectively? A Delphi Study of Germany's Non‐Formal Education Sector
by: Kim Alina Lüdtke, et al.
Published: (2024)
by: Kim Alina Lüdtke, et al.
Published: (2024)
Staircase Streaming for Low-Latency Multi-Agent Inference
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Watching the FedWatch
by: Stefano Bonini, et al.
Published: (2025)
by: Stefano Bonini, et al.
Published: (2025)
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
by: Akl, James, et al.
Published: (2026)
by: Akl, James, et al.
Published: (2026)
How to Use Research Evidence Well in Education
by: Rickinson, Mark, et al.
Published: (2025)
by: Rickinson, Mark, et al.
Published: (2025)
How Well Can Differential Privacy Be Audited in One Run?
by: Keinan, Amit, et al.
Published: (2025)
by: Keinan, Amit, et al.
Published: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
by: Giannou, Angeliki, et al.
Published: (2024)
by: Giannou, Angeliki, et al.
Published: (2024)
How Can Generative AI Enhance the Well-being of Blind?
by: Bendel, Oliver
Published: (2024)
by: Bendel, Oliver
Published: (2024)
Living Well with a Disability: How Libraries Can Help.
by: Klauber, Julie
Published: (1998)
by: Klauber, Julie
Published: (1998)
El momento de la verdad = Crossing the bridge [VHS - Video]
by: Mike Binder
by: Mike Binder
Similar Items
-
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
by: Thapa, Rahul, et al.
Published: (2024) -
Data Diversification Methods In Alignment Enhance Math Performance In LLMs
by: Dokmeci, Berkan, et al.
Published: (2025) -
EchoNet-Quality: Denoising Echocardiograms via Deep Generative Modeling of Ultrasound Noise
by: Choi, David, et al.
Published: (2025) -
Automated Interpretable 2D Video Extraction from 3D Echocardiography
by: Vukadinovic, Milos, et al.
Published: (2025) -
Disentangling Reasoning and Knowledge in Medical Large Language Models
by: Thapa, Rahul, et al.
Published: (2025)