MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Feilong, Liu, Yijiang, Huang, Yi, Wang, Hao, Tian, Miren, Yu, Ya-Qi, Liao, Minghui, Wu, Jihao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pulmonary Tuberculosis Edge Diagnosis System Based on MindSpore Framework: Low-cost and High-precision Implementation with Ascend 310 Chip
by: Li, HaoYu
Published: (2025)
by: Li, HaoYu
Published: (2025)
Classification based deep learning models for lung cancer and disease using medical images
by: Chaddad, Ahmad, et al.
Published: (2025)
by: Chaddad, Ahmad, et al.
Published: (2025)
Just Noticeable Difference for Large Multimodal Models
by: Chen, Zijian, et al.
Published: (2025)
by: Chen, Zijian, et al.
Published: (2025)
Towards Effective and Efficient Context-aware Nucleus Detection in Histopathology Whole Slide Images
by: Shui, Zhongyi, et al.
Published: (2025)
by: Shui, Zhongyi, et al.
Published: (2025)
ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model
by: Kondo, Satoshi
Published: (2025)
by: Kondo, Satoshi
Published: (2025)
Joint Learning Neuronal Skeleton and Brain Circuit Topology with Permutation Invariant Encoders for Neuron Classification
by: Liao, Minghui, et al.
Published: (2023)
by: Liao, Minghui, et al.
Published: (2023)
MEMO: Dataset and Methods for Robust Multimodal Retinal Image Registration with Large or Small Vessel Density Differences
by: Wang, Chiao-Yi, et al.
Published: (2023)
by: Wang, Chiao-Yi, et al.
Published: (2023)
From Images to Point Clouds: An Efficient Solution for Cross-media Blind Quality Assessment without Annotated Training
by: Liu, Yipeng, et al.
Published: (2025)
by: Liu, Yipeng, et al.
Published: (2025)
NeRFCodec: Neural Feature Compression Meets Neural Radiance Fields for Memory-Efficient Scene Representation
by: Li, Sicheng, et al.
Published: (2024)
by: Li, Sicheng, et al.
Published: (2024)
Stable Optimization for Large Vision Model Based Deep Image Prior in Cone-Beam CT Reconstruction
by: Wu, Minghui, et al.
Published: (2022)
by: Wu, Minghui, et al.
Published: (2022)
DSCENet: Dynamic Screening and Clinical-Enhanced Multimodal Fusion for MPNs Subtype Classification
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
Restore-RWKV: Efficient and Effective Medical Image Restoration with RWKV
by: Yang, Zhiwen, et al.
Published: (2024)
by: Yang, Zhiwen, et al.
Published: (2024)
Simplifying Multimodality: Unimodal Approach to Multimodal Challenges in Radiology with General-Domain Large Language Model
by: Cho, Seonhee, et al.
Published: (2024)
by: Cho, Seonhee, et al.
Published: (2024)
Boosting Convolution with Efficient MLP-Permutation for Volumetric Medical Image Segmentation
by: Lin, Yi, et al.
Published: (2023)
by: Lin, Yi, et al.
Published: (2023)
Towards Consistent Object Detection via LiDAR-Camera Synergy
by: Luo, Kai, et al.
Published: (2024)
by: Luo, Kai, et al.
Published: (2024)
M3-CVC: Controllable Video Compression with Multimodal Generative Models
by: Wan, Rui, et al.
Published: (2024)
by: Wan, Rui, et al.
Published: (2024)
EI-Nexus: Towards Unmediated and Flexible Inter-Modality Local Feature Extraction and Matching for Event-Image Data
by: Yi, Zhonghua, et al.
Published: (2024)
by: Yi, Zhonghua, et al.
Published: (2024)
Cohort-Individual Cooperative Learning for Multimodal Cancer Survival Analysis
by: Zhou, Huajun, et al.
Published: (2024)
by: Zhou, Huajun, et al.
Published: (2024)
Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms
by: Stojanovski, David, et al.
Published: (2024)
by: Stojanovski, David, et al.
Published: (2024)
Deep Ensembling with Multimodal Image Fusion for Efficient Classification of Lung Cancer
by: Pal, Surochita, et al.
Published: (2025)
by: Pal, Surochita, et al.
Published: (2025)
R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
by: Li, Chunyi, et al.
Published: (2024)
by: Li, Chunyi, et al.
Published: (2024)
MoME: Mixture of Multimodal Experts for Cancer Survival Prediction
by: Xiong, Conghao, et al.
Published: (2024)
by: Xiong, Conghao, et al.
Published: (2024)
Toward Efficient Deep Blind RAW Image Restoration
by: Conde, Marcos V., et al.
Published: (2024)
by: Conde, Marcos V., et al.
Published: (2024)
HESSO: Towards Automatic Efficient and User Friendly Any Neural Network Training and Pruning
by: Chen, Tianyi, et al.
Published: (2024)
by: Chen, Tianyi, et al.
Published: (2024)
Towards General Text-guided Image Synthesis for Customized Multimodal Brain MRI Generation
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
Towards Scalable and Robust White Matter Lesion Localization via Multimodal Deep Learning
by: Machnio, Julia, et al.
Published: (2025)
by: Machnio, Julia, et al.
Published: (2025)
GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
by: Chen, Pengcheng, et al.
Published: (2024)
by: Chen, Pengcheng, et al.
Published: (2024)
Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data
by: Zhou, Yiming, et al.
Published: (2026)
by: Zhou, Yiming, et al.
Published: (2026)
MS2Edge: Towards Energy-Efficient and Crisp Edge Detection with Multi-Scale Residual Learning in SNNs
by: Fan, Yimeng, et al.
Published: (2025)
by: Fan, Yimeng, et al.
Published: (2025)
Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare
by: Zhu, Hanwei, et al.
Published: (2024)
by: Zhu, Hanwei, et al.
Published: (2024)
Copy-Move Detection in Optical Microscopy: A Segmentation Network and A Dataset
by: Shao, Hao-Chiang, et al.
Published: (2024)
by: Shao, Hao-Chiang, et al.
Published: (2024)
Efficient Quality Control of Whole Slide Pathology Images with Human-in-the-loop Training
by: Patil, Abhijeet, et al.
Published: (2024)
by: Patil, Abhijeet, et al.
Published: (2024)
OmniLens: Towards Universal Lens Aberration Correction via LensLib-to-Specific Domain Adaptation
by: Jiang, Qi, et al.
Published: (2024)
by: Jiang, Qi, et al.
Published: (2024)
NAFRSSR: a Lightweight Recursive Network for Efficient Stereo Image Super-Resolution
by: Chen, Yihong, et al.
Published: (2024)
by: Chen, Yihong, et al.
Published: (2024)
Cross-Scan Mamba with Masked Training for Robust Spectral Imaging
by: Tian, Wenzhe, et al.
Published: (2024)
by: Tian, Wenzhe, et al.
Published: (2024)
Towards Defining an Efficient and Expandable File Format for AI-Generated Contents
by: Gao, Yixin, et al.
Published: (2024)
by: Gao, Yixin, et al.
Published: (2024)
Towards Generalizable Tumor Synthesis
by: Chen, Qi, et al.
Published: (2024)
by: Chen, Qi, et al.
Published: (2024)
Diffusion Model Driven Test-Time Image Adaptation for Robust Skin Lesion Classification
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
Evaluation of Geolocation Capabilities of Multimodal Large Language Models and Analysis of Associated Privacy Risks
by: Zhang, Xian, et al.
Published: (2025)
by: Zhang, Xian, et al.
Published: (2025)
MTT-Bench: Predicting Social Dominance in Mice via Multimodal Large Language Models
by: Chen, Yunquan, et al.
Published: (2026)
by: Chen, Yunquan, et al.
Published: (2026)
Similar Items
-
Pulmonary Tuberculosis Edge Diagnosis System Based on MindSpore Framework: Low-cost and High-precision Implementation with Ascend 310 Chip
by: Li, HaoYu
Published: (2025) -
Classification based deep learning models for lung cancer and disease using medical images
by: Chaddad, Ahmad, et al.
Published: (2025) -
Just Noticeable Difference for Large Multimodal Models
by: Chen, Zijian, et al.
Published: (2025) -
Towards Effective and Efficient Context-aware Nucleus Detection in Histopathology Whole Slide Images
by: Shui, Zhongyi, et al.
Published: (2025) -
ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model
by: Kondo, Satoshi
Published: (2025)