PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Tianyi, Shan, Rong, Wu, Junjie, Huang, Jiadeng, Wang, Teng, Zhu, Jiachen, Chen, Wenteng, Tu, Minxin, Dou, Quantao, Wang, Zhaoxiang, Zhang, Changwang, Zhang, Weinan, Wang, Jun, Lin, Jianghao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OSCAR: Optimization-Steered Agentic Planning for Composed Image Retrieval
by: Wang, Teng, et al.
Published: (2026)
by: Wang, Teng, et al.
Published: (2026)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026)
by: Zhou, Qianrui, et al.
Published: (2026)
Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2023)
by: Zhou, Qianrui, et al.
Published: (2023)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
Personality-Enhanced Multimodal Depression Detection in the Elderly
by: Wang, Honghong, et al.
Published: (2025)
by: Wang, Honghong, et al.
Published: (2025)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2025)
by: Zhou, Qianrui, et al.
Published: (2025)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
by: Pu, Junfu, et al.
Published: (2026)
by: Pu, Junfu, et al.
Published: (2026)
Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning
by: Hu, Jiayun, et al.
Published: (2025)
by: Hu, Jiayun, et al.
Published: (2025)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
by: Liu, Shuo, et al.
Published: (2024)
by: Liu, Shuo, et al.
Published: (2024)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
by: Zhang, Yongjian, et al.
Published: (2025)
by: Zhang, Yongjian, et al.
Published: (2025)
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
by: Luo, Bingjun, et al.
Published: (2025)
by: Luo, Bingjun, et al.
Published: (2025)
Prototypical Prompting for Text-to-image Person Re-identification
by: Yan, Shuanglin, et al.
Published: (2024)
by: Yan, Shuanglin, et al.
Published: (2024)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
by: Xu, Pengju, et al.
Published: (2025)
by: Xu, Pengju, et al.
Published: (2025)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
by: Deng, Weihui, et al.
Published: (2024)
by: Deng, Weihui, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
by: Peng, Yuezhang, et al.
Published: (2025)
by: Peng, Yuezhang, et al.
Published: (2025)
Enhancing Image-Text Matching with Adaptive Feature Aggregation
by: Wang, Zuhui, et al.
Published: (2024)
by: Wang, Zuhui, et al.
Published: (2024)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
by: Gao, Lancheng, et al.
Published: (2025)
by: Gao, Lancheng, et al.
Published: (2025)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
by: Huang, Yiheng, et al.
Published: (2024)
by: Huang, Yiheng, et al.
Published: (2024)
Towards Structure-aware Model for Multi-modal Knowledge Graph Completion
by: Li, Linyu, et al.
Published: (2025)
by: Li, Linyu, et al.
Published: (2025)
DREAM: A Dual Representation Learning Model for Multimodal Recommendation
by: Zhang, Kangning, et al.
Published: (2024)
by: Zhang, Kangning, et al.
Published: (2024)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
by: Zhang, Wang, et al.
Published: (2024)
by: Zhang, Wang, et al.
Published: (2024)
U-Sticker: A Large-Scale Multi-Domain User Sticker Dataset for Retrieval and Personalization
by: Chee, Heng Er Metilda, et al.
Published: (2025)
by: Chee, Heng Er Metilda, et al.
Published: (2025)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance Capture
by: Jin, Yitong, et al.
Published: (2024)
by: Jin, Yitong, et al.
Published: (2024)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
by: Meng, Jiahao, et al.
Published: (2026)
by: Meng, Jiahao, et al.
Published: (2026)
Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
by: Zhang, Beibei, et al.
Published: (2025)
by: Zhang, Beibei, et al.
Published: (2025)
Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
by: Wang, Sen, et al.
Published: (2024)
by: Wang, Sen, et al.
Published: (2024)
MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios
by: Li, Xiangxian, et al.
Published: (2025)
by: Li, Xiangxian, et al.
Published: (2025)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
by: Zhang, Xuling, et al.
Published: (2024)
by: Zhang, Xuling, et al.
Published: (2024)
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
by: Hu, Huanran, et al.
Published: (2026)
by: Hu, Huanran, et al.
Published: (2026)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
by: Wong, Kam Kwai, et al.
Published: (2023)
by: Wong, Kam Kwai, et al.
Published: (2023)
DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
by: Zhao, Kangran, et al.
Published: (2025)
by: Zhao, Kangran, et al.
Published: (2025)
Perceptual-oriented Learned Image Compression with Dynamic Kernel
by: Fu, Nianxiang, et al.
Published: (2024)
by: Fu, Nianxiang, et al.
Published: (2024)
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
by: Wang, Xiaohan, et al.
Published: (2026)
by: Wang, Xiaohan, et al.
Published: (2026)
Similar Items
-
OSCAR: Optimization-Steered Agentic Planning for Composed Image Retrieval
by: Wang, Teng, et al.
Published: (2026) -
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026) -
Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition
by: Wang, Yifan, et al.
Published: (2026) -
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2023) -
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)