KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goswami, Parthaw, Deep, Jaynto Goswami |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation
von: Yeafi, Ashfak, et al.
Veröffentlicht: (2026)
von: Yeafi, Ashfak, et al.
Veröffentlicht: (2026)
L-MCAT: Unpaired Multimodal Transformer with Contrastive Attention for Label-Efficient Satellite Image Classification
von: Goswami, Mitul, et al.
Veröffentlicht: (2025)
von: Goswami, Mitul, et al.
Veröffentlicht: (2025)
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Neuroplastic Modular Framework: Cross-Domain Image Classification of Garbage and Industrial Surfaces
von: Ghosh, Debojyoti, et al.
Veröffentlicht: (2025)
von: Ghosh, Debojyoti, et al.
Veröffentlicht: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)
von: Wu, Yin, et al.
Veröffentlicht: (2025)
Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation
von: Zhao, Ran, et al.
Veröffentlicht: (2026)
von: Zhao, Ran, et al.
Veröffentlicht: (2026)
KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
von: Huang, Hsin-Ping, et al.
Veröffentlicht: (2024)
von: Huang, Hsin-Ping, et al.
Veröffentlicht: (2024)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
von: Choi, Changin, et al.
Veröffentlicht: (2025)
von: Choi, Changin, et al.
Veröffentlicht: (2025)
KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
von: Ma, Xinyu, et al.
Veröffentlicht: (2025)
HybridSOMSpikeNet: A Deep Model with Differentiable Soft Self-Organizing Maps and Spiking Dynamics for Waste Classification
von: Ghosh, Debojyoti, et al.
Veröffentlicht: (2025)
von: Ghosh, Debojyoti, et al.
Veröffentlicht: (2025)
Convolutional Neural Network Segmentation for Satellite Imagery Data to Identify Landforms Using U-Net Architecture
von: Goswami, Mitul, et al.
Veröffentlicht: (2025)
von: Goswami, Mitul, et al.
Veröffentlicht: (2025)
Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
ACIL: Active Class Incremental Learning for Image Classification
von: Bhattacharya, Aditya R., et al.
Veröffentlicht: (2026)
von: Bhattacharya, Aditya R., et al.
Veröffentlicht: (2026)
Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
von: Lompo, Boammani Aser, et al.
Veröffentlicht: (2025)
von: Lompo, Boammani Aser, et al.
Veröffentlicht: (2025)
Object Retrieval for Visual Question Answering with Outside Knowledge
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
von: Goswami, Prajnan, et al.
Veröffentlicht: (2026)
von: Goswami, Prajnan, et al.
Veröffentlicht: (2026)
Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency
von: Xue, Xi, et al.
Veröffentlicht: (2025)
von: Xue, Xi, et al.
Veröffentlicht: (2025)
Calibrating Higher-Order Statistics for Few-Shot Class-Incremental Learning with Pre-trained Vision Transformers
von: Goswami, Dipam, et al.
Veröffentlicht: (2024)
von: Goswami, Dipam, et al.
Veröffentlicht: (2024)
CapHDR2IR: Caption-Driven Transfer from Visible Light to Infrared Domain
von: Peng, Jingchao, et al.
Veröffentlicht: (2024)
von: Peng, Jingchao, et al.
Veröffentlicht: (2024)
MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
von: Zhu, Minghao, et al.
Veröffentlicht: (2024)
von: Zhu, Minghao, et al.
Veröffentlicht: (2024)
Composed Image Retrieval for Training-Free Domain Conversion
von: Efthymiadis, Nikos, et al.
Veröffentlicht: (2024)
von: Efthymiadis, Nikos, et al.
Veröffentlicht: (2024)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
von: Mi, Li, et al.
Veröffentlicht: (2024)
von: Mi, Li, et al.
Veröffentlicht: (2024)
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
von: Tan, Cheng, et al.
Veröffentlicht: (2024)
von: Tan, Cheng, et al.
Veröffentlicht: (2024)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
Route, Retrieve, Reflect, Repair: Self-Improving Agentic Framework for Visual Detection and Linguistic Reasoning in Medical Imaging
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2026)
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2026)
VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
von: Kang, Hyeonsu, et al.
Veröffentlicht: (2025)
von: Kang, Hyeonsu, et al.
Veröffentlicht: (2025)
Additive Manufacturing Processes Protocol Prediction by Artificial Intelligence using X-ray Computed Tomography data
von: Khod, Sunita, et al.
Veröffentlicht: (2025)
von: Khod, Sunita, et al.
Veröffentlicht: (2025)
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
von: Dey, Sainath, et al.
Veröffentlicht: (2025)
von: Dey, Sainath, et al.
Veröffentlicht: (2025)
Psychological stress during Examination and its estimation by handwriting in answer script
von: Kumar, Abhijeet, et al.
Veröffentlicht: (2025)
von: Kumar, Abhijeet, et al.
Veröffentlicht: (2025)
FlashMix: Fast Map-Free LiDAR Localization via Feature Mixing and Contrastive-Constrained Accelerated Training
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
von: Jamil, Sofia, et al.
Veröffentlicht: (2025)
Hierarchical Matching and Reasoning for Multi-Query Image Retrieval
von: Ji, Zhong, et al.
Veröffentlicht: (2023)
von: Ji, Zhong, et al.
Veröffentlicht: (2023)
Knowledge-Intensive Video Generation
von: Wang, Chenxu, et al.
Veröffentlicht: (2026)
von: Wang, Chenxu, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval
von: Wang, Ziwei, et al.
Veröffentlicht: (2024)
von: Wang, Ziwei, et al.
Veröffentlicht: (2024)
Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
von: Li, Jun, et al.
Veröffentlicht: (2025)
von: Li, Jun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation
von: Yeafi, Ashfak, et al.
Veröffentlicht: (2026) -
L-MCAT: Unpaired Multimodal Transformer with Contrastive Attention for Label-Efficient Satellite Image Classification
von: Goswami, Mitul, et al.
Veröffentlicht: (2025) -
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
von: Wu, Di, et al.
Veröffentlicht: (2025) -
Neuroplastic Modular Framework: Cross-Domain Image Classification of Garbage and Industrial Surfaces
von: Ghosh, Debojyoti, et al.
Veröffentlicht: (2025) -
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)