MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Fan, Dong, Xingping, Yu, Xin, Luo, Wenhan, Liu, Wei, Zhang, Kaihao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
von: Yang, Rui, et al.
Veröffentlicht: (2024)
von: Yang, Rui, et al.
Veröffentlicht: (2024)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution
von: Zhao, Hanli, et al.
Veröffentlicht: (2026)
von: Zhao, Hanli, et al.
Veröffentlicht: (2026)
Learning Efficient Unsupervised Satellite Image-based Building Damage Detection
von: Zhang, Yiyun, et al.
Veröffentlicht: (2023)
von: Zhang, Yiyun, et al.
Veröffentlicht: (2023)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
von: Kalakonda, Sai Shashank, et al.
Veröffentlicht: (2024)
von: Kalakonda, Sai Shashank, et al.
Veröffentlicht: (2024)
TMFNet: Two-Stream Multi-Channels Fusion Networks for Color Image Operation Chain Detection
von: Niu, Yakun, et al.
Veröffentlicht: (2024)
von: Niu, Yakun, et al.
Veröffentlicht: (2024)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
von: Zhu, Haodong, et al.
Veröffentlicht: (2025)
von: Zhu, Haodong, et al.
Veröffentlicht: (2025)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
von: Zhang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Zhang, Wenfeng, et al.
Veröffentlicht: (2026)
Depth and Image Fusion for Road Obstacle Detection Using Stereo Camera
von: Perezyabov, Oleg, et al.
Veröffentlicht: (2025)
von: Perezyabov, Oleg, et al.
Veröffentlicht: (2025)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
von: Wang, Xue, et al.
Veröffentlicht: (2026)
von: Wang, Xue, et al.
Veröffentlicht: (2026)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Fine-grained Image Retrieval via Dual-Vision Adaptation
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
Advancing Unsupervised Low-light Image Enhancement: Noise Estimation, Illumination Interpolation, and Self-Regulation
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
von: Li, Linfei, et al.
Veröffentlicht: (2025)
von: Li, Linfei, et al.
Veröffentlicht: (2025)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
PRVR: Partially Relevant Video Retrieval
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
von: Sun, Qianqian, et al.
Veröffentlicht: (2025)
von: Sun, Qianqian, et al.
Veröffentlicht: (2025)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Opinion-Unaware Blind Image Quality Assessment using Multi-Scale Deep Feature Statistics
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision
von: Yin, Kangsheng, et al.
Veröffentlicht: (2025)
von: Yin, Kangsheng, et al.
Veröffentlicht: (2025)
DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
von: Li, Xinzhu, et al.
Veröffentlicht: (2025)
von: Li, Xinzhu, et al.
Veröffentlicht: (2025)
Pseudo-triplet Guided Few-shot Composed Image Retrieval
von: Hou, Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Bohan, et al.
Veröffentlicht: (2024)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
von: Zhang, Long, et al.
Veröffentlicht: (2025)
von: Zhang, Long, et al.
Veröffentlicht: (2025)
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
von: Dang, Yunkai, et al.
Veröffentlicht: (2025)
von: Dang, Yunkai, et al.
Veröffentlicht: (2025)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
Hyperspectral Image Fusion with Spectral-Band and Fusion-Scale Agnosticism
von: Liang, Yu-Jie, et al.
Veröffentlicht: (2026)
von: Liang, Yu-Jie, et al.
Veröffentlicht: (2026)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
von: Zhu, Lingsi, et al.
Veröffentlicht: (2026)
von: Zhu, Lingsi, et al.
Veröffentlicht: (2026)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
von: Yao, Ting, et al.
Veröffentlicht: (2024)
von: Yao, Ting, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
von: Yang, Rui, et al.
Veröffentlicht: (2024) -
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
von: Zhu, Yixin, et al.
Veröffentlicht: (2026) -
EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution
von: Zhao, Hanli, et al.
Veröffentlicht: (2026) -
Learning Efficient Unsupervised Satellite Image-based Building Damage Detection
von: Zhang, Yiyun, et al.
Veröffentlicht: (2023) -
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)