SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Ren-Di, Lin, Yu-Yen, Yang, Huei-Fang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and Similarity
di: Wu, Ren-Di, et al.
Pubblicazione: (2024)
di: Wu, Ren-Di, et al.
Pubblicazione: (2024)
Does CLIP perceive art the same way we do?
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
di: Jin, Hang, et al.
Pubblicazione: (2025)
di: Jin, Hang, et al.
Pubblicazione: (2025)
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
di: Louison, Nikita, et al.
Pubblicazione: (2024)
di: Louison, Nikita, et al.
Pubblicazione: (2024)
Multimodal Structure-Aware Quantum Data Processing
di: Hawashin, Hala, et al.
Pubblicazione: (2024)
di: Hawashin, Hala, et al.
Pubblicazione: (2024)
PDFMathTranslate: Scientific Document Translation Preserving Layouts
di: Ouyang, Rongxin, et al.
Pubblicazione: (2025)
di: Ouyang, Rongxin, et al.
Pubblicazione: (2025)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
di: Rashid, Muhammad, et al.
Pubblicazione: (2026)
di: Rashid, Muhammad, et al.
Pubblicazione: (2026)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
di: Viveiros, André G., et al.
Pubblicazione: (2025)
di: Viveiros, André G., et al.
Pubblicazione: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
di: Cai, Weibin, et al.
Pubblicazione: (2025)
di: Cai, Weibin, et al.
Pubblicazione: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
di: Zhang, Junwen, et al.
Pubblicazione: (2025)
di: Zhang, Junwen, et al.
Pubblicazione: (2025)
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
di: Gad, Eyad, et al.
Pubblicazione: (2025)
di: Gad, Eyad, et al.
Pubblicazione: (2025)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2026)
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2026)
Unraveling Media Perspectives: A Comprehensive Methodology Combining Large Language Models, Topic Modeling, Sentiment Analysis, and Ontology Learning to Analyse Media Bias
di: Jähde, Orlando, et al.
Pubblicazione: (2025)
di: Jähde, Orlando, et al.
Pubblicazione: (2025)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
di: Komurcu, Kursat, et al.
Pubblicazione: (2026)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
di: Kang, Xueyang, et al.
Pubblicazione: (2026)
di: Kang, Xueyang, et al.
Pubblicazione: (2026)
Dual-sensing driving detection model
di: K, Leon C. C., et al.
Pubblicazione: (2025)
di: K, Leon C. C., et al.
Pubblicazione: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
di: Svystun, Serhii, et al.
Pubblicazione: (2025)
di: Svystun, Serhii, et al.
Pubblicazione: (2025)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
di: Cui, Hejie, et al.
Pubblicazione: (2024)
di: Cui, Hejie, et al.
Pubblicazione: (2024)
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
di: Zhang, Shengkai, et al.
Pubblicazione: (2024)
di: Zhang, Shengkai, et al.
Pubblicazione: (2024)
Image-based Facial Rig Inversion
di: Yang, Tianxiang, et al.
Pubblicazione: (2025)
di: Yang, Tianxiang, et al.
Pubblicazione: (2025)
A Cost-Effective Eye-Tracker for Early Detection of Mild Cognitive Impairment
di: Greco, Danilo, et al.
Pubblicazione: (2024)
di: Greco, Danilo, et al.
Pubblicazione: (2024)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
di: Kambhatla, Akhila, et al.
Pubblicazione: (2025)
di: Kambhatla, Akhila, et al.
Pubblicazione: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
di: Baydemir, Poyraz
Pubblicazione: (2025)
di: Baydemir, Poyraz
Pubblicazione: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
di: Siddiqui, Yousuf Ahmed, et al.
Pubblicazione: (2025)
di: Siddiqui, Yousuf Ahmed, et al.
Pubblicazione: (2025)
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
di: Rao, Penghao, et al.
Pubblicazione: (2025)
di: Rao, Penghao, et al.
Pubblicazione: (2025)
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
di: Rakib, Mazharul Islam, et al.
Pubblicazione: (2025)
di: Rakib, Mazharul Islam, et al.
Pubblicazione: (2025)
Isolated Sign Language Recognition with Segmentation and Pose Estimation
di: Perkins, Daniel, et al.
Pubblicazione: (2025)
di: Perkins, Daniel, et al.
Pubblicazione: (2025)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
di: Papyan, Narek, et al.
Pubblicazione: (2024)
di: Papyan, Narek, et al.
Pubblicazione: (2024)
Data Augmentation and Resolution Enhancement using GANs and Diffusion Models for Tree Segmentation
di: Ferreira, Alessandro dos Santos, et al.
Pubblicazione: (2025)
di: Ferreira, Alessandro dos Santos, et al.
Pubblicazione: (2025)
Real Time Human Detection by Unmanned Aerial Vehicles
di: Guettala, Walid, et al.
Pubblicazione: (2024)
di: Guettala, Walid, et al.
Pubblicazione: (2024)
A Landmark-Aware Visual Navigation Dataset
di: Johnson, Faith, et al.
Pubblicazione: (2024)
di: Johnson, Faith, et al.
Pubblicazione: (2024)
Benchmarking Vision Language Models on German Factual Data
di: Peinl, René, et al.
Pubblicazione: (2025)
di: Peinl, René, et al.
Pubblicazione: (2025)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2024)
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2024)
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
di: Qi, Dekang, et al.
Pubblicazione: (2026)
di: Qi, Dekang, et al.
Pubblicazione: (2026)
LLM-supported document separation for printed reviews from zbMATH Open
di: Pluzhnikov, Ivan, et al.
Pubblicazione: (2026)
di: Pluzhnikov, Ivan, et al.
Pubblicazione: (2026)
IDOL: Instant Photorealistic 3D Human Creation from a Single Image
di: Zhuang, Yiyu, et al.
Pubblicazione: (2024)
di: Zhuang, Yiyu, et al.
Pubblicazione: (2024)
Force-Aware 3D Contact Modeling for Stable Grasp Generation
di: Chen, Zhuo, et al.
Pubblicazione: (2025)
di: Chen, Zhuo, et al.
Pubblicazione: (2025)
Banana Ripeness Level Classification using a Simple CNN Model Trained with Real and Synthetic Datasets
di: Chuquimarca, Luis, et al.
Pubblicazione: (2025)
di: Chuquimarca, Luis, et al.
Pubblicazione: (2025)
Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
di: Nasriddinov, Firdavs, et al.
Pubblicazione: (2025)
di: Nasriddinov, Firdavs, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and Similarity
di: Wu, Ren-Di, et al.
Pubblicazione: (2024) -
Does CLIP perceive art the same way we do?
di: Asperti, Andrea, et al.
Pubblicazione: (2025) -
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
di: Jin, Hang, et al.
Pubblicazione: (2025) -
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
di: Louison, Nikita, et al.
Pubblicazione: (2024) -
Multimodal Structure-Aware Quantum Data Processing
di: Hawashin, Hala, et al.
Pubblicazione: (2024)