Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Young Kyun, Kim, Donghyun, Meng, Zihang, Huynh, Dat, Lim, Ser-Nam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Composing Object Relations and Attributes for Image-Text Matching
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
von: Han, Gyuwon, et al.
Veröffentlicht: (2026)
von: Han, Gyuwon, et al.
Veröffentlicht: (2026)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
von: He, Bo, et al.
Veröffentlicht: (2024)
von: He, Bo, et al.
Veröffentlicht: (2024)
Towards Chunk-Wise Generation for Long Videos
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
Composed Multi-modal Retrieval: A Survey of Approaches and Applications
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
Semi-supervised Medical Image Segmentation via Geometry-aware Consistency Training
von: Liu, Zihang, et al.
Veröffentlicht: (2022)
von: Liu, Zihang, et al.
Veröffentlicht: (2022)
VideoMerge: Towards Training-free Long Video Generation
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
von: Gao, Bo, et al.
Veröffentlicht: (2026)
von: Gao, Bo, et al.
Veröffentlicht: (2026)
CoLLM: A Large Language Model for Composed Image Retrieval
von: Huynh, Chuong, et al.
Veröffentlicht: (2025)
von: Huynh, Chuong, et al.
Veröffentlicht: (2025)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
Prompt Augmentation for Self-supervised Text-guided Image Manipulation
von: Bodur, Rumeysa, et al.
Veröffentlicht: (2024)
von: Bodur, Rumeysa, et al.
Veröffentlicht: (2024)
MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval
von: Ge, Xuri, et al.
Veröffentlicht: (2026)
von: Ge, Xuri, et al.
Veröffentlicht: (2026)
FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval
von: Park, Jeong-Woo, et al.
Veröffentlicht: (2025)
von: Park, Jeong-Woo, et al.
Veröffentlicht: (2025)
Metric Compatible Training for Online Backfilling in Large-Scale Retrieval
von: Seo, Seonguk, et al.
Veröffentlicht: (2023)
von: Seo, Seonguk, et al.
Veröffentlicht: (2023)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
von: Ko, Dohwan, et al.
Veröffentlicht: (2025)
von: Ko, Dohwan, et al.
Veröffentlicht: (2025)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
von: Liu, Zheyuan, et al.
Veröffentlicht: (2023)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2023)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
MGHanD: Multi-modal Guidance for authentic Hand Diffusion
von: Eum, Taehyeon, et al.
Veröffentlicht: (2025)
von: Eum, Taehyeon, et al.
Veröffentlicht: (2025)
HOMIE: Histopathology Omni-modal Embedding for Pathology Composed Retrieval
von: Zhou, Qifeng, et al.
Veröffentlicht: (2025)
von: Zhou, Qifeng, et al.
Veröffentlicht: (2025)
Towards Semi-supervised Dual-modal Semantic Segmentation
von: Dong, Qiulei, et al.
Veröffentlicht: (2024)
von: Dong, Qiulei, et al.
Veröffentlicht: (2024)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
Dual Relation Alignment for Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2023)
von: Jiang, Xintong, et al.
Veröffentlicht: (2023)
AirSketch: Generative Motion to Sketch
von: Lim, Hui Xian Grace, et al.
Veröffentlicht: (2024)
von: Lim, Hui Xian Grace, et al.
Veröffentlicht: (2024)
FSViewFusion: Few-Shots View Generation of Novel Objects
von: Hussain, Rukhshanda, et al.
Veröffentlicht: (2024)
von: Hussain, Rukhshanda, et al.
Veröffentlicht: (2024)
RA-SGG: Retrieval-Augmented Scene Graph Generation Framework via Multi-Prototype Learning
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2024)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
von: Kim, Hyunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyunsoo, et al.
Veröffentlicht: (2025)
Arbitrary-Scale Image Generation and Upsampling using Latent Diffusion Model and Implicit Neural Decoder
von: Kim, Jinseok, et al.
Veröffentlicht: (2024)
von: Kim, Jinseok, et al.
Veröffentlicht: (2024)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
von: Chen, Zining, et al.
Veröffentlicht: (2025)
von: Chen, Zining, et al.
Veröffentlicht: (2025)
Instance-Level Composed Image Retrieval
von: Psomas, Bill, et al.
Veröffentlicht: (2025)
von: Psomas, Bill, et al.
Veröffentlicht: (2025)
Zero Shot Composed Image Retrieval
von: Kakarla, Santhosh, et al.
Veröffentlicht: (2025)
von: Kakarla, Santhosh, et al.
Veröffentlicht: (2025)
Composed Image Retrieval for Remote Sensing
von: Psomas, Bill, et al.
Veröffentlicht: (2024)
von: Psomas, Bill, et al.
Veröffentlicht: (2024)
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
GOAL: Global-local Object Alignment Learning
von: Choi, Hyungyu, et al.
Veröffentlicht: (2025)
von: Choi, Hyungyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024) -
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024) -
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024) -
Composing Object Relations and Attributes for Image-Text Matching
von: Pham, Khoi, et al.
Veröffentlicht: (2024) -
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
von: Han, Gyuwon, et al.
Veröffentlicht: (2026)