VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Bufang, He, Lixing, Liu, Kaiwei, Yan, Zhenyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Hallucinations of Multi-modal Large Language Models
by: Huang, Wen, et al.
Published: (2024)
by: Huang, Wen, et al.
Published: (2024)
An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
by: He, Hulingxiao, et al.
Published: (2025)
by: He, Hulingxiao, et al.
Published: (2025)
PAVE: Patching and Adapting Video Large Language Models
by: Liu, Zhuoming, et al.
Published: (2025)
by: Liu, Zhuoming, et al.
Published: (2025)
Refusing Safe Prompts for Multi-modal Large Language Models
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
by: Lin, Zichuan, et al.
Published: (2025)
by: Lin, Zichuan, et al.
Published: (2025)
Multi-modal Vision Pre-training for Medical Image Analysis
by: Rui, Shaohao, et al.
Published: (2024)
by: Rui, Shaohao, et al.
Published: (2024)
PolyPath: Adapting a Large Multimodal Model for Multi-slide Pathology Report Generation
by: Ahmed, Faruk, et al.
Published: (2025)
by: Ahmed, Faruk, et al.
Published: (2025)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
by: Liu, Dongyang, et al.
Published: (2024)
by: Liu, Dongyang, et al.
Published: (2024)
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
by: Dong, Hao, et al.
Published: (2025)
by: Dong, Hao, et al.
Published: (2025)
Adapting Vision-Language Models for Evaluating World Models
by: Hendriksen, Mariya, et al.
Published: (2025)
by: Hendriksen, Mariya, et al.
Published: (2025)
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
by: Li, Shengzhi, et al.
Published: (2024)
by: Li, Shengzhi, et al.
Published: (2024)
Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
by: Cao, Xusheng, et al.
Published: (2025)
by: Cao, Xusheng, et al.
Published: (2025)
Generative Multi-modal Models are Good Class-Incremental Learners
by: Cao, Xusheng, et al.
Published: (2024)
by: Cao, Xusheng, et al.
Published: (2024)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
by: Sun, Hai-Long, et al.
Published: (2025)
by: Sun, Hai-Long, et al.
Published: (2025)
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
by: Liu, Zhongye, et al.
Published: (2024)
by: Liu, Zhongye, et al.
Published: (2024)
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
FEAT: A Multi-Agent Forensic AI System with Domain-Adapted Large Language Model for Automated Cause-of-Death Analysis
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
by: Wang, Yubo, et al.
Published: (2024)
by: Wang, Yubo, et al.
Published: (2024)
Adapting Large Multimodal Models to Distribution Shifts: The Role of In-Context Learning
by: Zhou, Guanglin, et al.
Published: (2024)
by: Zhou, Guanglin, et al.
Published: (2024)
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
by: Dong, Yuhao, et al.
Published: (2026)
by: Dong, Yuhao, et al.
Published: (2026)
Visual Question Decomposition on Multimodal Large Language Models
by: Zhang, Haowei, et al.
Published: (2024)
by: Zhang, Haowei, et al.
Published: (2024)
GrInAdapt: Scaling Retinal Vessel Structural Map Segmentation Through Grounding, Integrating and Adapting Multi-device, Multi-site, and Multi-modal Fundus Domains
by: Liu, Zixuan, et al.
Published: (2025)
by: Liu, Zixuan, et al.
Published: (2025)
Exploring the Transferability of Visual Prompting for Multimodal Large Language Models
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
by: Yoon, Heegeon, et al.
Published: (2026)
by: Yoon, Heegeon, et al.
Published: (2026)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
by: Lian, Chenyu, et al.
Published: (2025)
by: Lian, Chenyu, et al.
Published: (2025)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
by: Li, Zhuowei, et al.
Published: (2025)
by: Li, Zhuowei, et al.
Published: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
Adapting Segment Anything Model to Melanoma Segmentation in Microscopy Slide Images
by: Liu, Qingyuan, et al.
Published: (2024)
by: Liu, Qingyuan, et al.
Published: (2024)
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
by: Chen, Haokun, et al.
Published: (2025)
by: Chen, Haokun, et al.
Published: (2025)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
by: Kang, Weitai, et al.
Published: (2025)
by: Kang, Weitai, et al.
Published: (2025)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design
by: Liu, Anjie, et al.
Published: (2026)
by: Liu, Anjie, et al.
Published: (2026)
Visual Language Model based Cross-modal Semantic Communication Systems
by: Jiang, Feibo, et al.
Published: (2024)
by: Jiang, Feibo, et al.
Published: (2024)
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
by: Miao, Yanting, et al.
Published: (2026)
by: Miao, Yanting, et al.
Published: (2026)
Similar Items
-
Visual Hallucinations of Multi-modal Large Language Models
by: Huang, Wen, et al.
Published: (2024) -
An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
by: Jiang, Siyang, et al.
Published: (2025) -
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
by: He, Hulingxiao, et al.
Published: (2025) -
PAVE: Patching and Adapting Video Large Language Models
by: Liu, Zhuoming, et al.
Published: (2025) -
Refusing Safe Prompts for Multi-modal Large Language Models
by: Shao, Zedian, et al.
Published: (2024)