AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xue, Junxiao, Deng, Quan, Hu, Tingqi, Si, Meicong, Yin, Xinyi, Shi, Yunyun, Wu, Xuecheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MooD: Perception-Enhanced Efficient Affective Image Editing via Continuous Valence-Arousal Modeling
by: Yin, Xinyi, et al.
Published: (2026)
by: Yin, Xinyi, et al.
Published: (2026)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024)
by: Xue, Junxiao, et al.
Published: (2024)
AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control
by: Chen, Shi, et al.
Published: (2026)
by: Chen, Shi, et al.
Published: (2026)
Disentangling Hardness from Noise: An Uncertainty-Driven Model-Agnostic Framework for Long-Tailed Remote Sensing Classification
by: Ding, Chi, et al.
Published: (2026)
by: Ding, Chi, et al.
Published: (2026)
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
by: Su, Tiancheng, et al.
Published: (2025)
by: Su, Tiancheng, et al.
Published: (2025)
3A-YOLO: New Real-Time Object Detectors with Triple Discriminative Awareness and Coordinated Representations
by: Wu, Xuecheng, et al.
Published: (2024)
by: Wu, Xuecheng, et al.
Published: (2024)
A Trustworthy Method for Multimodal Emotion Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
by: Zhang, Zhengxuan, et al.
Published: (2025)
by: Zhang, Zhengxuan, et al.
Published: (2025)
TaSR-RAG: Taxonomy-guided Structured Reasoning for Retrieval-Augmented Generation
by: Sun, Jiashuo, et al.
Published: (2026)
by: Sun, Jiashuo, et al.
Published: (2026)
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues
by: Xue, Liang, et al.
Published: (2025)
by: Xue, Liang, et al.
Published: (2025)
AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
by: Wu, Ruipu, et al.
Published: (2025)
by: Wu, Ruipu, et al.
Published: (2025)
Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation
by: Wang, Baisen, et al.
Published: (2024)
by: Wang, Baisen, et al.
Published: (2024)
FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
by: Li, Gen, et al.
Published: (2026)
by: Li, Gen, et al.
Published: (2026)
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
Affective Video Content Analysis: Decade Review and New Perspectives
by: Xue, Junxiao, et al.
Published: (2023)
by: Xue, Junxiao, et al.
Published: (2023)
AirSpatialBot: A Spatially-Aware Aerial Agent for Fine-Grained Vehicle Attribute Recognization and Retrieval
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
by: Sun, Yubo, et al.
Published: (2025)
by: Sun, Yubo, et al.
Published: (2025)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
by: Zeng, Nianbo, et al.
Published: (2025)
by: Zeng, Nianbo, et al.
Published: (2025)
Mitigating Hallucination in Financial Retrieval-Augmented Generation via Fine-Grained Knowledge Verification
by: Yin, Taoye, et al.
Published: (2026)
by: Yin, Taoye, et al.
Published: (2026)
FT-RAG: A Fine-grained Retrieval-Augmented Generation Framework for Complex Table Reasoning
by: Guo, Zebin, et al.
Published: (2026)
by: Guo, Zebin, et al.
Published: (2026)
GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
by: Duan, Xinyi, et al.
Published: (2026)
by: Duan, Xinyi, et al.
Published: (2026)
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing
by: Xue, Fengjian, et al.
Published: (2026)
by: Xue, Fengjian, et al.
Published: (2026)
AeroScene: Progressive Scene Synthesis for Aerial Robotics
by: Vu, Nghia, et al.
Published: (2026)
by: Vu, Nghia, et al.
Published: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
GrepRAG: An Empirical Study and Optimization of Grep-Like Retrieval for Code Completion
by: Wang, Baoyi, et al.
Published: (2026)
by: Wang, Baoyi, et al.
Published: (2026)
CardioRAG: A Retrieval-Augmented Generation Framework for Multimodal Chagas Disease Detection
by: Shen, Zhengyang, et al.
Published: (2025)
by: Shen, Zhengyang, et al.
Published: (2025)
LLM-Driven Discovery of High-Entropy Catalysts via Retrieval-Augmented Generation
by: Scientists, AI, et al.
Published: (2026)
by: Scientists, AI, et al.
Published: (2026)
CrisiSense-RAG: Crisis Sensing Multimodal Retrieval-Augmented Generation for Rapid Disaster Impact Assessment
by: Xiao, Yiming, et al.
Published: (2026)
by: Xiao, Yiming, et al.
Published: (2026)
FGTR: Fine-Grained Multi-Table Retrieval via Hierarchical LLM Reasoning
by: Sun, Chaojie, et al.
Published: (2026)
by: Sun, Chaojie, et al.
Published: (2026)
RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance
by: Mortaheb, Matin, et al.
Published: (2025)
by: Mortaheb, Matin, et al.
Published: (2025)
Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis
by: Chaudhari, Shravan, et al.
Published: (2025)
by: Chaudhari, Shravan, et al.
Published: (2025)
HASH-RAG: Bridging Deep Hashing with Retriever for Efficient, Fine Retrieval and Augmented Generation
by: Guo, Jinyu, et al.
Published: (2025)
by: Guo, Jinyu, et al.
Published: (2025)
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges
by: Dhole, Kaustubh D., et al.
Published: (2024)
by: Dhole, Kaustubh D., et al.
Published: (2024)
Boosting Conversational Question Answering with Fine-Grained Retrieval-Augmentation and Self-Check
by: Ye, Linhao, et al.
Published: (2024)
by: Ye, Linhao, et al.
Published: (2024)
Similar Items
-
MooD: Perception-Enhanced Efficient Affective Image Editing via Continuous Valence-Arousal Modeling
by: Yin, Xinyi, et al.
Published: (2026) -
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024) -
AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control
by: Chen, Shi, et al.
Published: (2026) -
Disentangling Hardness from Noise: An Uncertainty-Driven Model-Agnostic Framework for Long-Tailed Remote Sensing Classification
by: Ding, Chi, et al.
Published: (2026) -
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
by: Wu, Xuecheng, et al.
Published: (2025)