Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Lyou, Eunyi, Lee, Doyeon, Kim, Jooeun, Lee, Joonseok |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)
by: Lee, Junho, et al.
Published: (2026)
Relation-Aware Meta-Learning for Zero-shot Sketch-Based Image Retrieval
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Dynamic Multi-level Weighted Alignment Network for Zero-shot Sketch-based Image Retrieval
by: Su, Hanwen, et al.
Published: (2025)
by: Su, Hanwen, et al.
Published: (2025)
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)
by: Hahm, Jaehoon, et al.
Published: (2024)
QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
by: Lee, Doyeon, et al.
Published: (2026)
by: Lee, Doyeon, et al.
Published: (2026)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
by: Lee, Inseo, et al.
Published: (2025)
by: Lee, Inseo, et al.
Published: (2025)
Multi-Granularity Representation Learning for Sketch-based Dynamic Face Image Retrieval
by: Wang, Liang, et al.
Published: (2023)
by: Wang, Liang, et al.
Published: (2023)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
by: Lee, Soeun, et al.
Published: (2024)
by: Lee, Soeun, et al.
Published: (2024)
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024)
by: Ha, Seongsu, et al.
Published: (2024)
Dual-Modal Prompting for Sketch-Based Image Retrieval
by: Gao, Liying, et al.
Published: (2024)
by: Gao, Liying, et al.
Published: (2024)
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
by: Li, Haiwen, et al.
Published: (2024)
by: Li, Haiwen, et al.
Published: (2024)
Freeview Sketching: View-Aware Fine-Grained Sketch-Based Image Retrieval
by: Sain, Aneeshan, et al.
Published: (2024)
by: Sain, Aneeshan, et al.
Published: (2024)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval
by: Du, Yongchao, et al.
Published: (2024)
by: Du, Yongchao, et al.
Published: (2024)
InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation
by: Kim, Chanran, et al.
Published: (2024)
by: Kim, Chanran, et al.
Published: (2024)
Camera Splatting for Continuous View Optimization
by: Lee, Gahye, et al.
Published: (2025)
by: Lee, Gahye, et al.
Published: (2025)
Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
by: Ryu, Suho, et al.
Published: (2025)
by: Ryu, Suho, et al.
Published: (2025)
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning
by: Singha, Mainak, et al.
Published: (2024)
by: Singha, Mainak, et al.
Published: (2024)
Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
by: Kwon, Soyeong, et al.
Published: (2024)
by: Kwon, Soyeong, et al.
Published: (2024)
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025)
by: Shin, Jeongwoo, et al.
Published: (2025)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
by: Chen, Zining, et al.
Published: (2025)
by: Chen, Zining, et al.
Published: (2025)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
by: Kang, Jeonghun, et al.
Published: (2025)
by: Kang, Jeonghun, et al.
Published: (2025)
Zero-shot Quantization: A Comprehensive Survey
by: Kim, Minjun, et al.
Published: (2025)
by: Kim, Minjun, et al.
Published: (2025)
SketchQL Demonstration: Zero-shot Video Moment Querying with Sketches
by: Wu, Renzhi, et al.
Published: (2024)
by: Wu, Renzhi, et al.
Published: (2024)
How to Handle Sketch-Abstraction in Sketch-Based Image Retrieval?
by: Koley, Subhadeep, et al.
Published: (2024)
by: Koley, Subhadeep, et al.
Published: (2024)
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
by: Gu, Geonmo, et al.
Published: (2023)
by: Gu, Geonmo, et al.
Published: (2023)
Flashback: Memory-Driven Zero-shot, Real-time Video Anomaly Detection
by: Lee, Hyogun, et al.
Published: (2025)
by: Lee, Hyogun, et al.
Published: (2025)
Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
by: Suo, Yucheng, et al.
Published: (2024)
by: Suo, Yucheng, et al.
Published: (2024)
PDV: Prompt Directional Vectors for Zero-shot Composed Image Retrieval
by: Tursun, Osman, et al.
Published: (2025)
by: Tursun, Osman, et al.
Published: (2025)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
by: Kim, Donghyeong, et al.
Published: (2025)
by: Kim, Donghyeong, et al.
Published: (2025)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026)
by: Kim, Sumin, et al.
Published: (2026)
Latent Representation Matters: Human-like Sketches in One-shot Drawing Tasks
by: Boutin, Victor, et al.
Published: (2024)
by: Boutin, Victor, et al.
Published: (2024)
Training-free Zero-shot Composed Image Retrieval with Local Concept Reranking
by: Sun, Shitong, et al.
Published: (2023)
by: Sun, Shitong, et al.
Published: (2023)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
Similar Items
-
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026) -
Relation-Aware Meta-Learning for Zero-shot Sketch-Based Image Retrieval
by: Liu, Yang, et al.
Published: (2024) -
Dynamic Multi-level Weighted Alignment Network for Zero-shot Sketch-based Image Retrieval
by: Su, Hanwen, et al.
Published: (2025) -
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025) -
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)