Visual RAG: Expanding MLLM visual knowledge without fine-tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Bonomo, Mirco, Bianco, Simone |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Uncertainty modeling for fine-tuned implicit functions
por: Susmelj, Anna, et al.
Publicado: (2024)
por: Susmelj, Anna, et al.
Publicado: (2024)
Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
por: Ahmed, Ammar, et al.
Publicado: (2025)
por: Ahmed, Ammar, et al.
Publicado: (2025)
What explains the success of cross-modal fine-tuning with ORCA?
por: García-de-Herreros, Paloma, et al.
Publicado: (2024)
por: García-de-Herreros, Paloma, et al.
Publicado: (2024)
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
por: Yang, Kai, et al.
Publicado: (2023)
por: Yang, Kai, et al.
Publicado: (2023)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
por: Lin, Han, et al.
Publicado: (2025)
por: Lin, Han, et al.
Publicado: (2025)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
por: Li, Yi, et al.
Publicado: (2026)
por: Li, Yi, et al.
Publicado: (2026)
Transporting Task Vectors across Different Architectures without Training
por: Rinaldi, Filippo, et al.
Publicado: (2026)
por: Rinaldi, Filippo, et al.
Publicado: (2026)
Multimodal RAG Enhanced Visual Description
por: Jaiswal, Amit Kumar, et al.
Publicado: (2025)
por: Jaiswal, Amit Kumar, et al.
Publicado: (2025)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
por: Duan, Chengqi, et al.
Publicado: (2025)
por: Duan, Chengqi, et al.
Publicado: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
por: Wang, Yiheng, et al.
Publicado: (2026)
por: Wang, Yiheng, et al.
Publicado: (2026)
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
por: Wu, Diankun, et al.
Publicado: (2025)
por: Wu, Diankun, et al.
Publicado: (2025)
ALARM: Automated MLLM-Based Anomaly Detection in Complex-EnviRonment Monitoring with Uncertainty Quantification
por: Zhang, Congjing, et al.
Publicado: (2025)
por: Zhang, Congjing, et al.
Publicado: (2025)
Gamified crowd-sourcing of high-quality data for visual fine-tuning
por: Yadav, Shashank, et al.
Publicado: (2024)
por: Yadav, Shashank, et al.
Publicado: (2024)
Transferring Visual Explainability of Self-Explaining Models to Prediction-Only Models without Additional Training
por: Yoshikawa, Yuya, et al.
Publicado: (2025)
por: Yoshikawa, Yuya, et al.
Publicado: (2025)
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding
por: Bianchi, Lorenzo, et al.
Publicado: (2023)
por: Bianchi, Lorenzo, et al.
Publicado: (2023)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
por: Celona, Luigi, et al.
Publicado: (2023)
por: Celona, Luigi, et al.
Publicado: (2023)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
por: Ai, Jiaxin, et al.
Publicado: (2025)
por: Ai, Jiaxin, et al.
Publicado: (2025)
Enhancing Worldwide Image Geolocation by Ensembling Satellite-Based Ground-Level Attribute Predictors
por: Bianco, Michael J., et al.
Publicado: (2024)
por: Bianco, Michael J., et al.
Publicado: (2024)
Pre-Trained Model Recommendation for Downstream Fine-tuning
por: Bai, Jiameng, et al.
Publicado: (2024)
por: Bai, Jiameng, et al.
Publicado: (2024)
LoCO: Low-rank Compositional Rotation Fine-tuning
por: Nguyen, An, et al.
Publicado: (2026)
por: Nguyen, An, et al.
Publicado: (2026)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
por: Kim, Hyunjae, et al.
Publicado: (2024)
por: Kim, Hyunjae, et al.
Publicado: (2024)
Structured Unrestricted-Rank Matrices for Parameter Efficient Fine-tuning
por: Sehanobish, Arijit, et al.
Publicado: (2024)
por: Sehanobish, Arijit, et al.
Publicado: (2024)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
por: Hu, Xueyu, et al.
Publicado: (2025)
por: Hu, Xueyu, et al.
Publicado: (2025)
MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
por: Liu, Qinghua, et al.
Publicado: (2025)
por: Liu, Qinghua, et al.
Publicado: (2025)
Generative Dataset Distillation Based on Self-knowledge Distillation
por: Li, Longzhen, et al.
Publicado: (2025)
por: Li, Longzhen, et al.
Publicado: (2025)
Improving fine-grained understanding in image-text pre-training
por: Bica, Ioana, et al.
Publicado: (2024)
por: Bica, Ioana, et al.
Publicado: (2024)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
por: Shahzad, Sahibzada Adil, et al.
Publicado: (2026)
por: Shahzad, Sahibzada Adil, et al.
Publicado: (2026)
Adaptive Random Feature Regularization on Fine-tuning Deep Neural Networks
por: Yamaguchi, Shin'ya, et al.
Publicado: (2024)
por: Yamaguchi, Shin'ya, et al.
Publicado: (2024)
See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition
por: Si, Chongjie, et al.
Publicado: (2024)
por: Si, Chongjie, et al.
Publicado: (2024)
Video Reasoning without Training
por: Sridhar, Deepak, et al.
Publicado: (2025)
por: Sridhar, Deepak, et al.
Publicado: (2025)
Towards flexible perception with visual memory
por: Geirhos, Robert, et al.
Publicado: (2024)
por: Geirhos, Robert, et al.
Publicado: (2024)
GCAM: Gaussian and causal-attention model of food fine-grained recognition
por: Zhuang, Guohang, et al.
Publicado: (2024)
por: Zhuang, Guohang, et al.
Publicado: (2024)
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
por: Wu, Junda, et al.
Publicado: (2025)
por: Wu, Junda, et al.
Publicado: (2025)
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
por: Wang, Zhengbo, et al.
Publicado: (2024)
por: Wang, Zhengbo, et al.
Publicado: (2024)
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
por: Zhang, Gengwei, et al.
Publicado: (2024)
por: Zhang, Gengwei, et al.
Publicado: (2024)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
por: Metzen, Jan Hendrik, et al.
Publicado: (2023)
por: Metzen, Jan Hendrik, et al.
Publicado: (2023)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
por: Xu, Binqian, et al.
Publicado: (2024)
por: Xu, Binqian, et al.
Publicado: (2024)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
por: Wang, Chenxi, et al.
Publicado: (2024)
por: Wang, Chenxi, et al.
Publicado: (2024)
A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning
por: Gupta, Shashank, et al.
Publicado: (2025)
por: Gupta, Shashank, et al.
Publicado: (2025)
EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction
por: Lin, Hsi-Che, et al.
Publicado: (2025)
por: Lin, Hsi-Che, et al.
Publicado: (2025)
Ejemplares similares
-
Uncertainty modeling for fine-tuned implicit functions
por: Susmelj, Anna, et al.
Publicado: (2024) -
Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
por: Ahmed, Ammar, et al.
Publicado: (2025) -
What explains the success of cross-modal fine-tuning with ORCA?
por: García-de-Herreros, Paloma, et al.
Publicado: (2024) -
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
por: Yang, Kai, et al.
Publicado: (2023) -
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
por: Lin, Han, et al.
Publicado: (2025)