Transformer-empowered Multi-modal Item Embedding for Enhanced Image Search in E-Commerce
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Chang, Hou, Peng, Zeng, Anxiang, Yu, Han |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
von: Wang, Teng, et al.
Veröffentlicht: (2024)
von: Wang, Teng, et al.
Veröffentlicht: (2024)
Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items
von: Lin, Jianghao, et al.
Veröffentlicht: (2025)
von: Lin, Jianghao, et al.
Veröffentlicht: (2025)
Enhancing Incomplete Multi-modal Brain Tumor Segmentation with Intra-modal Asymmetry and Inter-modal Dependency
von: Liu, Weide, et al.
Veröffentlicht: (2024)
von: Liu, Weide, et al.
Veröffentlicht: (2024)
UAE: Universal Anatomical Embedding on Multi-modality Medical Images
von: Bai, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Bai, Xiaoyu, et al.
Veröffentlicht: (2023)
CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
Multi-modal Learnable Queries for Image Aesthetics Assessment
von: Xiong, Zhiwei, et al.
Veröffentlicht: (2024)
von: Xiong, Zhiwei, et al.
Veröffentlicht: (2024)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
ITEm: Unsupervised Image-Text Embedding Learning for eCommerce
von: Liao, Baohao, et al.
Veröffentlicht: (2023)
von: Liao, Baohao, et al.
Veröffentlicht: (2023)
CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2026)
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2026)
Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion
von: Li, Xilai, et al.
Veröffentlicht: (2023)
von: Li, Xilai, et al.
Veröffentlicht: (2023)
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
von: Wang, Jiepeng, et al.
Veröffentlicht: (2025)
von: Wang, Jiepeng, et al.
Veröffentlicht: (2025)
TransRef: Multi-Scale Reference Embedding Transformer for Reference-Guided Image Inpainting
von: Liu, Taorong, et al.
Veröffentlicht: (2023)
von: Liu, Taorong, et al.
Veröffentlicht: (2023)
HyCTAS: Multi-Objective Hybrid Convolution-Transformer Architecture Search for Real-Time Image Segmentation
von: Yu, Hongyuan, et al.
Veröffentlicht: (2024)
von: Yu, Hongyuan, et al.
Veröffentlicht: (2024)
Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset
von: You, Zhiyuan, et al.
Veröffentlicht: (2024)
von: You, Zhiyuan, et al.
Veröffentlicht: (2024)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
von: Liu, Tengfei, et al.
Veröffentlicht: (2024)
von: Liu, Tengfei, et al.
Veröffentlicht: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
von: Cui, Jialei, et al.
Veröffentlicht: (2025)
von: Cui, Jialei, et al.
Veröffentlicht: (2025)
FOLDER: Accelerating Multi-modal Large Language Models with Enhanced Performance
von: Wang, Haicheng, et al.
Veröffentlicht: (2025)
von: Wang, Haicheng, et al.
Veröffentlicht: (2025)
CLIP Multi-modal Hashing for Multimedia Retrieval
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
Shared Multi-modal Embedding Space for Face-Voice Association
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
Fine-grained Multi-class Nuclei Segmentation with Molecular-empowered All-in-SAM Model
von: Li, Xueyuan, et al.
Veröffentlicht: (2025)
von: Li, Xueyuan, et al.
Veröffentlicht: (2025)
MATCNN: Infrared and Visible Image Fusion Method Based on Multi-scale CNN with Attention Transformer
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
Enhancing Transformers Through Conditioned Embedded Tokens
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2025)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2025)
Training-Free Style Consistent Image Synthesis with Condition and Mask Guidance in E-Commerce
von: Li, Guandong
Veröffentlicht: (2024)
von: Li, Guandong
Veröffentlicht: (2024)
3D-aware Image Generation and Editing with Multi-modal Conditions
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
von: Yu, Bo, et al.
Veröffentlicht: (2025)
von: Yu, Bo, et al.
Veröffentlicht: (2025)
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
PhotoFramer: Multi-modal Image Composition Instruction
von: You, Zhiyuan, et al.
Veröffentlicht: (2025)
von: You, Zhiyuan, et al.
Veröffentlicht: (2025)
Modeling Multi-modal Cross-interaction for Multi-label Few-shot Image Classification Based on Local Feature Selection
von: Yan, Kun, et al.
Veröffentlicht: (2024)
von: Yan, Kun, et al.
Veröffentlicht: (2024)
PROFUSEme: PROstate Cancer Biochemical Recurrence Prediction via FUSEd Multi-modal Embeddings
von: You, Suhang, et al.
Veröffentlicht: (2025)
von: You, Suhang, et al.
Veröffentlicht: (2025)
VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs
von: Qian, Dongjun, et al.
Veröffentlicht: (2025)
von: Qian, Dongjun, et al.
Veröffentlicht: (2025)
MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2024)
von: Liu, Wenzhuo, et al.
Veröffentlicht: (2024)
CADFormer: Fine-Grained Cross-modal Alignment and Decoding Transformer for Referring Remote Sensing Image Segmentation
von: Liu, Maofu, et al.
Veröffentlicht: (2025)
von: Liu, Maofu, et al.
Veröffentlicht: (2025)
Advancing Re-Ranking with Multimodal Fusion and Target-Oriented Auxiliary Tasks in E-Commerce Search
von: Xu, Enqiang, et al.
Veröffentlicht: (2024)
von: Xu, Enqiang, et al.
Veröffentlicht: (2024)
Oracle Bone Inscriptions Multi-modal Dataset
von: Li, Bang, et al.
Veröffentlicht: (2024)
von: Li, Bang, et al.
Veröffentlicht: (2024)
RGB-T Object Detection via Group Shuffled Multi-receptive Attention and Multi-modal Supervision
von: Wang, Jinzhong, et al.
Veröffentlicht: (2024)
von: Wang, Jinzhong, et al.
Veröffentlicht: (2024)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
von: Huang, Minbin, et al.
Veröffentlicht: (2024)
von: Huang, Minbin, et al.
Veröffentlicht: (2024)
Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix
von: Liu, Peng
Veröffentlicht: (2021)
von: Liu, Peng
Veröffentlicht: (2021)
From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts
von: Li, Weiran, et al.
Veröffentlicht: (2025)
von: Li, Weiran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
von: Wang, Teng, et al.
Veröffentlicht: (2024) -
Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items
von: Lin, Jianghao, et al.
Veröffentlicht: (2025) -
Enhancing Incomplete Multi-modal Brain Tumor Segmentation with Intra-modal Asymmetry and Inter-modal Dependency
von: Liu, Weide, et al.
Veröffentlicht: (2024) -
UAE: Universal Anatomical Embedding on Multi-modality Medical Images
von: Bai, Xiaoyu, et al.
Veröffentlicht: (2023) -
CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)