Saved in:
| Main Authors: | Lozano, Alejandro, Sun, Min Woo, Burgess, James, Nirschl, Jeffrey J., Polzak, Christopher, Zhang, Yuhui, Chen, Liangyu, Gu, Jeffrey, Lopez, Ivan, Aklilu, Josiah, Rau, Anita, Katzer, Austin Wolfgang, Chiu, Collin, Zohar, Orr, Wang, Xiaohan, Song, Alfred Seunghoon, Chia-Chun, Chiang, Tibshirani, Robert, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.22727 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
by: Lozano, Alejandro, et al.
Published: (2025)
by: Lozano, Alejandro, et al.
Published: (2025)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
by: Chen, Liangyu, et al.
Published: (2025)
by: Chen, Liangyu, et al.
Published: (2025)
Revisiting Active Learning in the Era of Vision Foundation Models
by: Gupte, Sanket Rajan, et al.
Published: (2024)
by: Gupte, Sanket Rajan, et al.
Published: (2024)
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024)
by: Rau, Anita, et al.
Published: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024)
by: Aklilu, Josiah, et al.
Published: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
by: Lozano, Alejandro, et al.
Published: (2024)
by: Lozano, Alejandro, et al.
Published: (2024)
Temporal Preference Optimization for Long-Form Video Understanding
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
by: Rau, Anita, et al.
Published: (2025)
by: Rau, Anita, et al.
Published: (2025)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
Can Large Language Models Match the Conclusions of Systematic Reviews?
by: Polzak, Christopher, et al.
Published: (2025)
by: Polzak, Christopher, et al.
Published: (2025)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
CellFlux: Simulating Cellular Morphology Changes via Flow Matching
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes
by: Meleti, Uma, et al.
Published: (2026)
by: Meleti, Uma, et al.
Published: (2026)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
by: Sun, Min Woo, et al.
Published: (2025)
by: Sun, Min Woo, et al.
Published: (2025)
The Evaluation of Libraries: Considerations from a Research Perspective.
by: Katzer, Jeffrey
Published: (1977)
by: Katzer, Jeffrey
Published: (1977)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023)
by: Burgess, James, et al.
Published: (2023)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)
by: Endo, Mark, et al.
Published: (2024)
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
by: Tang, Bingda, et al.
Published: (2026)
by: Tang, Bingda, et al.
Published: (2026)
Closing the Modality Gap for Mixed Modality Search
by: Li, Binxu, et al.
Published: (2025)
by: Li, Binxu, et al.
Published: (2025)
NegVQA: Can Vision Language Models Understand Negation?
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Why Don't They Ask Questions?
by: Swope, Mary Jane, et al.
Published: (1972)
by: Swope, Mary Jane, et al.
Published: (1972)
Why are Visually-Grounded Language Models Bad at Image Classification?
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Tool Verification for Test-Time Reinforcement Learning
by: Liao, Ruotong, et al.
Published: (2026)
by: Liao, Ruotong, et al.
Published: (2026)
Three Forms of Stochastic Injection for Improved Distribution-to-Distribution Generative Modeling
by: Su, Shiye, et al.
Published: (2025)
by: Su, Shiye, et al.
Published: (2025)
Data or Language Supervision: What Makes CLIP Better than DINO?
by: Liu, Yiming, et al.
Published: (2025)
by: Liu, Yiming, et al.
Published: (2025)
CryoHype: Reconstructing a thousand cryo-EM structures with transformer-based hypernetworks
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
by: Shen, Xiaoxian, et al.
Published: (2026)
by: Shen, Xiaoxian, et al.
Published: (2026)
Integral Probability Metrics Meet Neural Networks: The Radon-Kolmogorov-Smirnov Test
by: Paik, Seunghoon, et al.
Published: (2023)
by: Paik, Seunghoon, et al.
Published: (2023)
Basic Inequalities for First-Order Optimization with Applications to Statistical Risk Analysis
by: Paik, Seunghoon, et al.
Published: (2025)
by: Paik, Seunghoon, et al.
Published: (2025)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)
by: Endo, Mark, et al.
Published: (2025)
Apollo: An Exploration of Video Understanding in Large Multimodal Models
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Describing Differences in Image Sets with Natural Language
by: Dunlap, Lisa, et al.
Published: (2023)
by: Dunlap, Lisa, et al.
Published: (2023)
Similar Items
-
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
by: Lozano, Alejandro, et al.
Published: (2025) -
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
by: Chen, Liangyu, et al.
Published: (2025) -
Revisiting Active Learning in the Era of Vision Foundation Models
by: Gupte, Sanket Rajan, et al.
Published: (2024) -
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024) -
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024)