Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and Similarity
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Ren-Di, Lin, Yu-Yen, Yang, Huei-Fang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
by: Wu, Ren-Di, et al.
Published: (2025)
by: Wu, Ren-Di, et al.
Published: (2025)
Multimodal Structure-Aware Quantum Data Processing
by: Hawashin, Hala, et al.
Published: (2024)
by: Hawashin, Hala, et al.
Published: (2024)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
by: Lentsch, Ted, et al.
Published: (2026)
by: Lentsch, Ted, et al.
Published: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
by: Lentsch, Ted, et al.
Published: (2024)
by: Lentsch, Ted, et al.
Published: (2024)
Subspace Clustering in Wavelet Packets Domain
by: Kopriva, Ivica, et al.
Published: (2024)
by: Kopriva, Ivica, et al.
Published: (2024)
LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Interpretable label-free self-guided subspace clustering
by: Kopriva, Ivica
Published: (2024)
by: Kopriva, Ivica
Published: (2024)
Extraction Of Cumulative Blobs From Dynamic Gestures
by: Naulakha, Rishabh, et al.
Published: (2025)
by: Naulakha, Rishabh, et al.
Published: (2025)
CAFCT-Net: A CNN-Transformer Hybrid Network with Contextual and Attentional Feature Fusion for Liver Tumor Segmentation
by: Kang, Ming, et al.
Published: (2024)
by: Kang, Ming, et al.
Published: (2024)
MultiFinRAG: An Optimized Multimodal Retrieval-Augmented Generation (RAG) Framework for Financial Question Answering
by: Gondhalekar, Chinmay, et al.
Published: (2025)
by: Gondhalekar, Chinmay, et al.
Published: (2025)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
by: Ko, Hyun-kyu, et al.
Published: (2024)
by: Ko, Hyun-kyu, et al.
Published: (2024)
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
by: Louison, Nikita, et al.
Published: (2024)
by: Louison, Nikita, et al.
Published: (2024)
Polygonizing Roof Segments from High-Resolution Aerial Images Using Yolov8-Based Edge Detection
by: Mei, Qipeng, et al.
Published: (2025)
by: Mei, Qipeng, et al.
Published: (2025)
WSCIF: A Weakly-Supervised Color Intelligence Framework for Tactical Anomaly Detection in Surveillance Keyframes
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
Floorplan2Guide: LLM-Guided Floorplan Parsing for BLV Indoor Navigation
by: Ayanzadeh, Aydin, et al.
Published: (2025)
by: Ayanzadeh, Aydin, et al.
Published: (2025)
A Multimodal Feature Distillation with Mamba-Transformer Network for Brain Tumor Segmentation with Incomplete Modalities
by: Kang, Ming, et al.
Published: (2024)
by: Kang, Ming, et al.
Published: (2024)
Dense Video Understanding with Gated Residual Tokenization
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
BGF-YOLO: Enhanced YOLOv8 with Multiscale Attentional Feature Fusion for Brain Tumor Detection
by: Kang, Ming, et al.
Published: (2023)
by: Kang, Ming, et al.
Published: (2023)
ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation
by: Livieris, Ioannis E., et al.
Published: (2026)
by: Livieris, Ioannis E., et al.
Published: (2026)
ASF-YOLO: A Novel YOLO Model with Attentional Scale Sequence Fusion for Cell Instance Segmentation
by: Kang, Ming, et al.
Published: (2023)
by: Kang, Ming, et al.
Published: (2023)
Neural Encoding for Image Recall: Human-Like Memory
by: Foussereau, Virgile, et al.
Published: (2024)
by: Foussereau, Virgile, et al.
Published: (2024)
From Particles to Agents: Hallucination as a Metric for Cognitive Friction in Spatial Simulation
by: Sánchez-Vaquerizo, Javier Argota, et al.
Published: (2026)
by: Sánchez-Vaquerizo, Javier Argota, et al.
Published: (2026)
A Practical Synthesis of Detecting AI-Generated Textual, Visual, and Audio Content
by: Cao, Lele
Published: (2025)
by: Cao, Lele
Published: (2025)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
by: Yang, Baoyao, et al.
Published: (2025)
by: Yang, Baoyao, et al.
Published: (2025)
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
by: Gad, Eyad, et al.
Published: (2025)
by: Gad, Eyad, et al.
Published: (2025)
The JPEG XL Image Coding System: History, Features, Coding Tools, Design Rationale, and Future
by: Sneyers, Jon, et al.
Published: (2025)
by: Sneyers, Jon, et al.
Published: (2025)
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
by: Gondhalekar, Chinmay, et al.
Published: (2025)
by: Gondhalekar, Chinmay, et al.
Published: (2025)
Structured Basis Function Networks: Loss-Centric Multi-Hypothesis Ensembles with Controllable Diversity
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
DeepShade: Enable Shade Simulation by Text-conditioned Image Generation
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
by: Kang, Xueyang, et al.
Published: (2026)
by: Kang, Xueyang, et al.
Published: (2026)
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
by: Kazemi, Amir, et al.
Published: (2024)
by: Kazemi, Amir, et al.
Published: (2024)
AI-Powered Augmented Reality for Satellite Assembly, Integration and Test
by: Patricio, Alvaro, et al.
Published: (2024)
by: Patricio, Alvaro, et al.
Published: (2024)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
by: Ayanzadeh, Aydin, et al.
Published: (2026)
by: Ayanzadeh, Aydin, et al.
Published: (2026)
Watermarking for AI Content Detection: A Review on Text, Visual, and Audio Modalities
by: Cao, Lele
Published: (2025)
by: Cao, Lele
Published: (2025)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
by: Kalušev, Vladimir, et al.
Published: (2025)
by: Kalušev, Vladimir, et al.
Published: (2025)
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
by: Jin, Hang, et al.
Published: (2025)
by: Jin, Hang, et al.
Published: (2025)
Trojan Detection Through Pattern Recognition for Large Language Models
by: Bhasin, Vedant, et al.
Published: (2025)
by: Bhasin, Vedant, et al.
Published: (2025)
CHORUS: An Agentic Framework for Generating Realistic Deliberation Data
by: Koursaris, A., et al.
Published: (2026)
by: Koursaris, A., et al.
Published: (2026)
CST-YOLO: A Novel Method for Blood Cell Detection Based on Improved YOLOv7 and CNN-Swin Transformer
by: Kang, Ming, et al.
Published: (2023)
by: Kang, Ming, et al.
Published: (2023)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
by: Kambhatla, Akhila, et al.
Published: (2025)
by: Kambhatla, Akhila, et al.
Published: (2025)
Similar Items
-
SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
by: Wu, Ren-Di, et al.
Published: (2025) -
Multimodal Structure-Aware Quantum Data Processing
by: Hawashin, Hala, et al.
Published: (2024) -
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
by: Lentsch, Ted, et al.
Published: (2026) -
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
by: Lentsch, Ted, et al.
Published: (2024) -
Subspace Clustering in Wavelet Packets Domain
by: Kopriva, Ivica, et al.
Published: (2024)