Saved in:
| Main Authors: | Mazor, Nir, Hope, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.17394 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
by: Lange, Bernard, et al.
Published: (2025)
by: Lange, Bernard, et al.
Published: (2025)
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024)
by: Ganz, Roy, et al.
Published: (2024)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
XDR-LVLM: An Explainable Vision-Language Large Model for Diabetic Retinopathy Diagnosis
by: Ito, Masato, et al.
Published: (2025)
by: Ito, Masato, et al.
Published: (2025)
Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation
by: Wu, Junde, et al.
Published: (2024)
by: Wu, Junde, et al.
Published: (2024)
MedM-VL: What Makes a Good Medical LVLM?
by: Shi, Yiming, et al.
Published: (2025)
by: Shi, Yiming, et al.
Published: (2025)
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
by: Garcia, Fernando Gabriela, et al.
Published: (2025)
by: Garcia, Fernando Gabriela, et al.
Published: (2025)
LVLM-Composer's Explicit Planning for Image Generation
by: Ramsey, Spencer, et al.
Published: (2025)
by: Ramsey, Spencer, et al.
Published: (2025)
RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
by: Chen, I-Hsiang, et al.
Published: (2026)
by: Chen, I-Hsiang, et al.
Published: (2026)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
by: Hsiao, Chi-Hsiang, et al.
Published: (2025)
by: Hsiao, Chi-Hsiang, et al.
Published: (2025)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
DuoGen: Towards General Purpose Interleaved Multimodal Generation
by: Shi, Min, et al.
Published: (2026)
by: Shi, Min, et al.
Published: (2026)
MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education
by: Benito, Miguel Diaz, et al.
Published: (2026)
by: Benito, Miguel Diaz, et al.
Published: (2026)
NICO-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Understanding the Nicotine Public Health Crisis
by: Serna-Aguilera, Manuel, et al.
Published: (2026)
by: Serna-Aguilera, Manuel, et al.
Published: (2026)
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
by: Zhu, Lanyun, et al.
Published: (2025)
by: Zhu, Lanyun, et al.
Published: (2025)
LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection
by: Liu, Qingyuan, et al.
Published: (2025)
by: Liu, Qingyuan, et al.
Published: (2025)
Large Language Model Aided Birt-Hogg-Dube Syndrome Diagnosis with Multimodal Retrieval-Augmented Generation
by: Li, Haoqing, et al.
Published: (2025)
by: Li, Haoqing, et al.
Published: (2025)
ASAP: Attention-Shift-Aware Pruning for Efficient LVLM Inference
by: Pathak, Surendra, et al.
Published: (2026)
by: Pathak, Surendra, et al.
Published: (2026)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
by: Park, Seongheon, et al.
Published: (2026)
by: Park, Seongheon, et al.
Published: (2026)
Uncovering Modality Discrepancy and Generalization Illusion for General-Purpose 3D Medical Segmentation
by: Zhang, Yichi, et al.
Published: (2026)
by: Zhang, Yichi, et al.
Published: (2026)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
by: Karim, A H M Rezaul, et al.
Published: (2025)
by: Karim, A H M Rezaul, et al.
Published: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models
by: Wang, Lehan, et al.
Published: (2025)
by: Wang, Lehan, et al.
Published: (2025)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
by: Qi, Jingyuan, et al.
Published: (2025)
by: Qi, Jingyuan, et al.
Published: (2025)
Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
by: Zhang, Wenchuan, et al.
Published: (2025)
by: Zhang, Wenchuan, et al.
Published: (2025)
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
by: Wu, Yuxuan, et al.
Published: (2025)
by: Wu, Yuxuan, et al.
Published: (2025)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
by: Wang, Qiuchen, et al.
Published: (2026)
by: Wang, Qiuchen, et al.
Published: (2026)
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
by: Sanguigni, Fulvio, et al.
Published: (2025)
by: Sanguigni, Fulvio, et al.
Published: (2025)
AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLM
by: Zhong, Li'an, et al.
Published: (2026)
by: Zhong, Li'an, et al.
Published: (2026)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
by: Zhu, Chenhui, et al.
Published: (2025)
by: Zhu, Chenhui, et al.
Published: (2025)
StreamingRAG: Real-time Contextual Retrieval and Generation Framework
by: Sankaradas, Murugan, et al.
Published: (2025)
by: Sankaradas, Murugan, et al.
Published: (2025)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
by: Cao, Linhan, et al.
Published: (2026)
by: Cao, Linhan, et al.
Published: (2026)
AstroRAG -- A Pagerank-Based Retrieval-Augmented Generation Pipeline for Question Answering in Astronomy
by: Wang, Zhifeng, et al.
Published: (2026)
by: Wang, Zhifeng, et al.
Published: (2026)
A Semantically-Aware Relevance Measure for Content-Based Medical Image Retrieval Evaluation
by: Wei, Xiaoyang, et al.
Published: (2025)
by: Wei, Xiaoyang, et al.
Published: (2025)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
by: Vo, Dinh-Khoi, et al.
Published: (2025)
by: Vo, Dinh-Khoi, et al.
Published: (2025)
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
by: Feng, Hengyi, et al.
Published: (2025)
by: Feng, Hengyi, et al.
Published: (2025)
LVLM-Aided Alignment of Task-Specific Vision Models
by: Koebler, Alexander, et al.
Published: (2025)
by: Koebler, Alexander, et al.
Published: (2025)
ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference
by: Huang, Zhaohong, et al.
Published: (2026)
by: Huang, Zhaohong, et al.
Published: (2026)
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
by: Xue, Junxiao, et al.
Published: (2026)
by: Xue, Junxiao, et al.
Published: (2026)
Similar Items
-
General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
by: Lange, Bernard, et al.
Published: (2025) -
Question Aware Vision Transformer for Multimodal Reasoning
by: Ganz, Roy, et al.
Published: (2024) -
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025) -
XDR-LVLM: An Explainable Vision-Language Large Model for Diabetic Retinopathy Diagnosis
by: Ito, Masato, et al.
Published: (2025) -
Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation
by: Wu, Junde, et al.
Published: (2024)