LaVy: Vietnamese Multimodal Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tran, Chi, Thanh, Huong Le |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring
von: Ho, Thanh Cong, et al.
Veröffentlicht: (2025)
von: Ho, Thanh Cong, et al.
Veröffentlicht: (2025)
Improving Multimodal Large Language Models Using Continual Learning
von: Srivastava, Shikhar, et al.
Veröffentlicht: (2024)
von: Srivastava, Shikhar, et al.
Veröffentlicht: (2024)
On Domain-Adaptive Post-Training for Multimodal Large Language Models
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
von: Yang, Xu, et al.
Veröffentlicht: (2023)
von: Yang, Xu, et al.
Veröffentlicht: (2023)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
Multimodal Latent Language Modeling with Next-Token Diffusion
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
von: Khangaonkar, Om, et al.
Veröffentlicht: (2026)
von: Khangaonkar, Om, et al.
Veröffentlicht: (2026)
Visual Question Decomposition on Multimodal Large Language Models
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
Jailbreaking Attack against Multimodal Large Language Model
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
von: Nguyen, Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2024)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
Towards Visual Text Grounding of Multimodal Large Language Model
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
von: Yue, Tongtian, et al.
Veröffentlicht: (2025)
von: Yue, Tongtian, et al.
Veröffentlicht: (2025)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
von: Ma, Chuofan, et al.
Veröffentlicht: (2024)
von: Ma, Chuofan, et al.
Veröffentlicht: (2024)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
von: Hong, Jindong, et al.
Veröffentlicht: (2025)
$ϕ$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2026)
von: Truong, Thanh-Dat, et al.
Veröffentlicht: (2026)
Text-Enhanced Data-free Approach for Federated Class-Incremental Learning
von: Tran, Minh-Tuan, et al.
Veröffentlicht: (2024)
von: Tran, Minh-Tuan, et al.
Veröffentlicht: (2024)
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
A Survey on Hallucination in Large Vision-Language Models
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models
von: Groot, Tobias, et al.
Veröffentlicht: (2024)
von: Groot, Tobias, et al.
Veröffentlicht: (2024)
MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
von: Ganesan, Mugilan, et al.
Veröffentlicht: (2025)
von: Ganesan, Mugilan, et al.
Veröffentlicht: (2025)
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2026)
von: Lau, Gregory Kang Ruey, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024) -
REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring
von: Ho, Thanh Cong, et al.
Veröffentlicht: (2025) -
Improving Multimodal Large Language Models Using Continual Learning
von: Srivastava, Shikhar, et al.
Veröffentlicht: (2024) -
On Domain-Adaptive Post-Training for Multimodal Large Language Models
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024) -
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
von: Chen, Feng, et al.
Veröffentlicht: (2024)