FITRep: Attention-Guided Item Representation via MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Guoxiao, Li, Ao, Qu, Tan, Xie, Qianlong, Wang, Xingxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918218869768192
author Zhang, Guoxiao
Li, Ao
Qu, Tan
Xie, Qianlong
Wang, Xingxing
author_facet Zhang, Guoxiao
Li, Ao
Qu, Tan
Xie, Qianlong
Wang, Xingxing
contents Online platforms usually suffer from user experience degradation due to near-duplicate items with similar visuals and text. While Multimodal Large Language Models (MLLMs) enable multimodal embedding, existing methods treat representations as black boxes, ignoring structural relationships (e.g., primary vs. auxiliary elements), leading to local structural collapse problem. To address this, inspired by Feature Integration Theory (FIT), we propose FITRep, the first attention-guided, white-box item representation framework for fine-grained item deduplication. FITRep consists of: (1) Concept Hierarchical Information Extraction (CHIE), using MLLMs to extract hierarchical semantic concepts; (2) Structure-Preserving Dimensionality Reduction (SPDR), an adaptive UMAP-based method for efficient information compression; and (3) FAISS-Based Clustering (FBC), a FAISS-based clustering that assigns each item a unique cluster id using FAISS. Deployed on Meituan's advertising system, FITRep achieves +3.60% CTR and +4.25% CPM gains in online A/B tests, demonstrating both effectiveness and real-world impact.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21389
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FITRep: Attention-Guided Item Representation via MLLMs
Zhang, Guoxiao
Li, Ao
Qu, Tan
Xie, Qianlong
Wang, Xingxing
Information Retrieval
Artificial Intelligence
Online platforms usually suffer from user experience degradation due to near-duplicate items with similar visuals and text. While Multimodal Large Language Models (MLLMs) enable multimodal embedding, existing methods treat representations as black boxes, ignoring structural relationships (e.g., primary vs. auxiliary elements), leading to local structural collapse problem. To address this, inspired by Feature Integration Theory (FIT), we propose FITRep, the first attention-guided, white-box item representation framework for fine-grained item deduplication. FITRep consists of: (1) Concept Hierarchical Information Extraction (CHIE), using MLLMs to extract hierarchical semantic concepts; (2) Structure-Preserving Dimensionality Reduction (SPDR), an adaptive UMAP-based method for efficient information compression; and (3) FAISS-Based Clustering (FBC), a FAISS-based clustering that assigns each item a unique cluster id using FAISS. Deployed on Meituan's advertising system, FITRep achieves +3.60% CTR and +4.25% CPM gains in online A/B tests, demonstrating both effectiveness and real-world impact.
title FITRep: Attention-Guided Item Representation via MLLMs
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2511.21389