Automatic Teaching Platform on Vision Language Retrieval Augmented Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gokhman, Ruslan, Li, Jialu, Zhang, Youshan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SparrowVQE: Visual Question Explanation for Course Content Understanding
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
von: Biyyala, Varun, et al.
Veröffentlicht: (2025)
von: Biyyala, Varun, et al.
Veröffentlicht: (2025)
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
von: Yu, Enhui, et al.
Veröffentlicht: (2026)
von: Yu, Enhui, et al.
Veröffentlicht: (2026)
FairRAG: Fair Human Generation via Fair Retrieval Augmentation
von: Shrestha, Robik, et al.
Veröffentlicht: (2024)
von: Shrestha, Robik, et al.
Veröffentlicht: (2024)
LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval
von: Zhang, Jian, et al.
Veröffentlicht: (2025)
von: Zhang, Jian, et al.
Veröffentlicht: (2025)
Confident Pseudo-labeled Diffusion Augmentation for Canine Cardiomegaly Detection
von: Zhang, Shiman, et al.
Veröffentlicht: (2025)
von: Zhang, Shiman, et al.
Veröffentlicht: (2025)
CrisiSense-RAG: Crisis Sensing Multimodal Retrieval-Augmented Generation for Rapid Disaster Impact Assessment
von: Xiao, Yiming, et al.
Veröffentlicht: (2026)
von: Xiao, Yiming, et al.
Veröffentlicht: (2026)
WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
von: Sultan, Rafi Ibn, et al.
Veröffentlicht: (2026)
von: Sultan, Rafi Ibn, et al.
Veröffentlicht: (2026)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
von: Lucy, Li, et al.
Veröffentlicht: (2026)
von: Lucy, Li, et al.
Veröffentlicht: (2026)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
von: Tan, Bryan Chen Zhengyu, et al.
Veröffentlicht: (2025)
von: Tan, Bryan Chen Zhengyu, et al.
Veröffentlicht: (2025)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
Unifying VLM-Guided Flow Matching and Spectral Anomaly Detection for Interpretable Veterinary Diagnosis
von: Wang, Pu, et al.
Veröffentlicht: (2026)
von: Wang, Pu, et al.
Veröffentlicht: (2026)
SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information
von: Sun, Jiashuo, et al.
Veröffentlicht: (2024)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
Urban Socio-Semantic Segmentation with Vision-Language Reasoning
von: Wang, Yu, et al.
Veröffentlicht: (2026)
von: Wang, Yu, et al.
Veröffentlicht: (2026)
RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
Stable Signer: Hierarchical Sign Language Generative Model
von: Fang, Sen, et al.
Veröffentlicht: (2025)
von: Fang, Sen, et al.
Veröffentlicht: (2025)
How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
von: Ming, Yifei, et al.
Veröffentlicht: (2023)
von: Ming, Yifei, et al.
Veröffentlicht: (2023)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspective of Creative Process
von: Bu, Yuyan, et al.
Veröffentlicht: (2024)
von: Bu, Yuyan, et al.
Veröffentlicht: (2024)
FairViT: Fair Vision Transformer via Adaptive Masking
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
von: Tian, Bowei, et al.
Veröffentlicht: (2024)
Identifying Implicit Social Biases in Vision-Language Models
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
von: Zhao, Yuming, et al.
Veröffentlicht: (2026)
von: Zhao, Yuming, et al.
Veröffentlicht: (2026)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
von: Fraser, Kathleen C., et al.
Veröffentlicht: (2024)
von: Fraser, Kathleen C., et al.
Veröffentlicht: (2024)
VQ-Jarvis: Retrieval-Augmented Video Restoration Agent with Sharp Vision and Fast Thought
von: Zhang, Xuanyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xuanyu, et al.
Veröffentlicht: (2026)
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes
von: Rivera, Antonio Carlos, et al.
Veröffentlicht: (2024)
von: Rivera, Antonio Carlos, et al.
Veröffentlicht: (2024)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Retrieval Augmented Comic Image Generation
von: Shui, Yunhao, et al.
Veröffentlicht: (2025)
von: Shui, Yunhao, et al.
Veröffentlicht: (2025)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
von: Zhang, Bo, et al.
Veröffentlicht: (2026)
Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models
von: Slyman, Eric, et al.
Veröffentlicht: (2024)
von: Slyman, Eric, et al.
Veröffentlicht: (2024)
Animated Territorial Data Extractor (ATDE): A Computer-Vision Method for Extracting Territorial Data from Animated Historical Maps
von: Alshamy, Hamza, et al.
Veröffentlicht: (2025)
von: Alshamy, Hamza, et al.
Veröffentlicht: (2025)
The Iconicity of the Generated Image
von: van Noord, Nanne, et al.
Veröffentlicht: (2025)
von: van Noord, Nanne, et al.
Veröffentlicht: (2025)
Generated Bias: Auditing Internal Bias Dynamics of Text-To-Image Generative Models
von: Mandal, Abhishek, et al.
Veröffentlicht: (2024)
von: Mandal, Abhishek, et al.
Veröffentlicht: (2024)
Vision-Language Models under Cultural and Inclusive Considerations
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2024)
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SparrowVQE: Visual Question Explanation for Course Content Understanding
von: Li, Jialu, et al.
Veröffentlicht: (2024) -
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
von: Biyyala, Varun, et al.
Veröffentlicht: (2025) -
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
von: Yu, Enhui, et al.
Veröffentlicht: (2026) -
FairRAG: Fair Human Generation via Fair Retrieval Augmentation
von: Shrestha, Robik, et al.
Veröffentlicht: (2024) -
LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval
von: Zhang, Jian, et al.
Veröffentlicht: (2025)