Visual RAG: Expanding MLLM visual knowledge without fine-tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Bonomo, Mirco, Bianco, Simone |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty modeling for fine-tuned implicit functions
by: Susmelj, Anna, et al.
Published: (2024)
by: Susmelj, Anna, et al.
Published: (2024)
Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
by: Yang, Kai, et al.
Published: (2023)
by: Yang, Kai, et al.
Published: (2023)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
by: Lin, Han, et al.
Published: (2025)
by: Lin, Han, et al.
Published: (2025)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
by: Li, Yi, et al.
Published: (2026)
by: Li, Yi, et al.
Published: (2026)
Transporting Task Vectors across Different Architectures without Training
by: Rinaldi, Filippo, et al.
Published: (2026)
by: Rinaldi, Filippo, et al.
Published: (2026)
Multimodal RAG Enhanced Visual Description
by: Jaiswal, Amit Kumar, et al.
Published: (2025)
by: Jaiswal, Amit Kumar, et al.
Published: (2025)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
by: Duan, Chengqi, et al.
Published: (2025)
by: Duan, Chengqi, et al.
Published: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
by: Wang, Yiheng, et al.
Published: (2026)
by: Wang, Yiheng, et al.
Published: (2026)
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
by: Wu, Diankun, et al.
Published: (2025)
by: Wu, Diankun, et al.
Published: (2025)
ALARM: Automated MLLM-Based Anomaly Detection in Complex-EnviRonment Monitoring with Uncertainty Quantification
by: Zhang, Congjing, et al.
Published: (2025)
by: Zhang, Congjing, et al.
Published: (2025)
Gamified crowd-sourcing of high-quality data for visual fine-tuning
by: Yadav, Shashank, et al.
Published: (2024)
by: Yadav, Shashank, et al.
Published: (2024)
Transferring Visual Explainability of Self-Explaining Models to Prediction-Only Models without Additional Training
by: Yoshikawa, Yuya, et al.
Published: (2025)
by: Yoshikawa, Yuya, et al.
Published: (2025)
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding
by: Bianchi, Lorenzo, et al.
Published: (2023)
by: Bianchi, Lorenzo, et al.
Published: (2023)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
by: Celona, Luigi, et al.
Published: (2023)
by: Celona, Luigi, et al.
Published: (2023)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
Enhancing Worldwide Image Geolocation by Ensembling Satellite-Based Ground-Level Attribute Predictors
by: Bianco, Michael J., et al.
Published: (2024)
by: Bianco, Michael J., et al.
Published: (2024)
Pre-Trained Model Recommendation for Downstream Fine-tuning
by: Bai, Jiameng, et al.
Published: (2024)
by: Bai, Jiameng, et al.
Published: (2024)
LoCO: Low-rank Compositional Rotation Fine-tuning
by: Nguyen, An, et al.
Published: (2026)
by: Nguyen, An, et al.
Published: (2026)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
by: Kim, Hyunjae, et al.
Published: (2024)
by: Kim, Hyunjae, et al.
Published: (2024)
Structured Unrestricted-Rank Matrices for Parameter Efficient Fine-tuning
by: Sehanobish, Arijit, et al.
Published: (2024)
by: Sehanobish, Arijit, et al.
Published: (2024)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
by: Hu, Xueyu, et al.
Published: (2025)
by: Hu, Xueyu, et al.
Published: (2025)
MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
by: Liu, Qinghua, et al.
Published: (2025)
by: Liu, Qinghua, et al.
Published: (2025)
Generative Dataset Distillation Based on Self-knowledge Distillation
by: Li, Longzhen, et al.
Published: (2025)
by: Li, Longzhen, et al.
Published: (2025)
Improving fine-grained understanding in image-text pre-training
by: Bica, Ioana, et al.
Published: (2024)
by: Bica, Ioana, et al.
Published: (2024)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
Adaptive Random Feature Regularization on Fine-tuning Deep Neural Networks
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition
by: Si, Chongjie, et al.
Published: (2024)
by: Si, Chongjie, et al.
Published: (2024)
Video Reasoning without Training
by: Sridhar, Deepak, et al.
Published: (2025)
by: Sridhar, Deepak, et al.
Published: (2025)
Towards flexible perception with visual memory
by: Geirhos, Robert, et al.
Published: (2024)
by: Geirhos, Robert, et al.
Published: (2024)
GCAM: Gaussian and causal-attention model of food fine-grained recognition
by: Zhuang, Guohang, et al.
Published: (2024)
by: Zhuang, Guohang, et al.
Published: (2024)
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
by: Wu, Junda, et al.
Published: (2025)
by: Wu, Junda, et al.
Published: (2025)
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
by: Zhang, Gengwei, et al.
Published: (2024)
by: Zhang, Gengwei, et al.
Published: (2024)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
by: Metzen, Jan Hendrik, et al.
Published: (2023)
by: Metzen, Jan Hendrik, et al.
Published: (2023)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning
by: Gupta, Shashank, et al.
Published: (2025)
by: Gupta, Shashank, et al.
Published: (2025)
EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction
by: Lin, Hsi-Che, et al.
Published: (2025)
by: Lin, Hsi-Che, et al.
Published: (2025)
Similar Items
-
Uncertainty modeling for fine-tuned implicit functions
by: Susmelj, Anna, et al.
Published: (2024) -
Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
by: Ahmed, Ammar, et al.
Published: (2025) -
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024) -
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
by: Yang, Kai, et al.
Published: (2023) -
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
by: Lin, Han, et al.
Published: (2025)