A Large-scale Medical Visual Task Adaptation Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Mo, Shentong, Luo, Xufang, Wang, Yansen, Li, Dongsheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
por: Mo, Shentong
Publicado: (2024)
por: Mo, Shentong
Publicado: (2024)
Improving Visual Representation Alignment Generation with GRPO
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
MultiMed: Massively Multimodal and Multitask Medical Understanding
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
por: Mo, Shentong
Publicado: (2024)
por: Mo, Shentong
Publicado: (2024)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
GMAIL: Generative Modality Alignment for generated Image Learning
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
por: Mo, Shentong, et al.
Publicado: (2023)
por: Mo, Shentong, et al.
Publicado: (2023)
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
por: Lin, Nie, et al.
Publicado: (2025)
por: Lin, Nie, et al.
Publicado: (2025)
IoT-LM: Large Multisensory Language Models for the Internet of Things
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration
por: Li, Zhili, et al.
Publicado: (2026)
por: Li, Zhili, et al.
Publicado: (2026)
Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE
por: Zhu, Xun, et al.
Publicado: (2024)
por: Zhu, Xun, et al.
Publicado: (2024)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
por: Mo, Shentong, et al.
Publicado: (2026)
por: Mo, Shentong, et al.
Publicado: (2026)
Unified Video-Language Pre-training with Synchronized Audio
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
Text-to-Audio Generation Synchronized with Videos
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
por: Mo, Shentong, et al.
Publicado: (2025)
por: Mo, Shentong, et al.
Publicado: (2025)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
por: Ming, Yifei, et al.
Publicado: (2024)
por: Ming, Yifei, et al.
Publicado: (2024)
Learning to Instruct for Visual Instruction Tuning
por: Zhou, Zhihan, et al.
Publicado: (2025)
por: Zhou, Zhihan, et al.
Publicado: (2025)
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
por: Hong, Jindong, et al.
Publicado: (2025)
por: Hong, Jindong, et al.
Publicado: (2025)
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification
por: Zhu, Xun, et al.
Publicado: (2026)
por: Zhu, Xun, et al.
Publicado: (2026)
CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset
por: Wang, Xiao, et al.
Publicado: (2024)
por: Wang, Xiao, et al.
Publicado: (2024)
AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models
por: Li, Xingjian, et al.
Publicado: (2025)
por: Li, Xingjian, et al.
Publicado: (2025)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
por: Pian, Weiguo, et al.
Publicado: (2024)
por: Pian, Weiguo, et al.
Publicado: (2024)
Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
por: Ahmadi, Mohammad Javad, et al.
Publicado: (2025)
por: Ahmadi, Mohammad Javad, et al.
Publicado: (2025)
MTLoRA: A Low-Rank Adaptation Approach for Efficient Multi-Task Learning
por: Agiza, Ahmed, et al.
Publicado: (2024)
por: Agiza, Ahmed, et al.
Publicado: (2024)
Relational Visual Similarity
por: Nguyen, Thao, et al.
Publicado: (2025)
por: Nguyen, Thao, et al.
Publicado: (2025)
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
por: Chen, Wenting, et al.
Publicado: (2025)
por: Chen, Wenting, et al.
Publicado: (2025)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
por: Woo, Byeongju, et al.
Publicado: (2026)
por: Woo, Byeongju, et al.
Publicado: (2026)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
por: Mai, Zheda, et al.
Publicado: (2025)
por: Mai, Zheda, et al.
Publicado: (2025)
MTL-MAD: Multi-Task Learners are Effective Medical Anomaly Detectors
por: Bercean, Bogdan Alexandru, et al.
Publicado: (2026)
por: Bercean, Bogdan Alexandru, et al.
Publicado: (2026)
Mars-Bench: A Benchmark for Evaluating Foundation Models for Mars Science Tasks
por: Purohit, Mirali, et al.
Publicado: (2025)
por: Purohit, Mirali, et al.
Publicado: (2025)
MAPSeg: Unified Unsupervised Domain Adaptation for Heterogeneous Medical Image Segmentation Based on 3D Masked Autoencoding and Pseudo-Labeling
por: Zhang, Xuzhe, et al.
Publicado: (2023)
por: Zhang, Xuzhe, et al.
Publicado: (2023)
Ejemplares similares
-
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
por: Mo, Shentong, et al.
Publicado: (2024) -
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
por: Mo, Shentong, et al.
Publicado: (2026) -
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
por: Mo, Shentong
Publicado: (2024) -
Improving Visual Representation Alignment Generation with GRPO
por: Mo, Shentong, et al.
Publicado: (2026) -
MultiMed: Massively Multimodal and Multitask Medical Understanding
por: Mo, Shentong, et al.
Publicado: (2024)