Visual Product Graph: Bridging Visual Products And Composite Images For End-to-End Style Recommendations
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Yue Li, Alexander, Ben, Antonenka, Mikhail, Mahadev, Rohan, Wu, Hao-yu, Kislyuk, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
by: Mahadev, Rohan, et al.
Published: (2026)
by: Mahadev, Rohan, et al.
Published: (2026)
V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
by: Guhan, Pooja, et al.
Published: (2025)
by: Guhan, Pooja, et al.
Published: (2025)
VSD-MOT: End-to-End Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Distillation
by: Du, Jun
Published: (2026)
by: Du, Jun
Published: (2026)
End-to-End Visual Autonomous Parking via Control-Aided Attention
by: Chen, Chao, et al.
Published: (2025)
by: Chen, Chao, et al.
Published: (2025)
PAVE: An End-to-End Dataset for Production Autonomous Vehicle Evaluation
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
by: Wu, Yuchen, et al.
Published: (2025)
by: Wu, Yuchen, et al.
Published: (2025)
StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving
by: Hao, Ruiyang, et al.
Published: (2025)
by: Hao, Ruiyang, et al.
Published: (2025)
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
by: Xie, Peijin, et al.
Published: (2025)
by: Xie, Peijin, et al.
Published: (2025)
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
by: Zheng, Peiru, et al.
Published: (2024)
by: Zheng, Peiru, et al.
Published: (2024)
Large-scale Reinforcement Learning for Diffusion Models
by: Zhang, Yinan, et al.
Published: (2024)
by: Zhang, Yinan, et al.
Published: (2024)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
by: Zhu, Wang, et al.
Published: (2023)
by: Zhu, Wang, et al.
Published: (2023)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
Pinterest Canvas: Large-Scale Image Generation at Pinterest
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
by: Salzmann, Tim, et al.
Published: (2024)
by: Salzmann, Tim, et al.
Published: (2024)
DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
Product of Experts for Visual Generation
by: Zhang, Yunzhi, et al.
Published: (2025)
by: Zhang, Yunzhi, et al.
Published: (2025)
Feature Corrective Transfer Learning: End-to-End Solutions to Object Detection in Non-Ideal Visual Conditions
by: Wei, Chuheng, et al.
Published: (2024)
by: Wei, Chuheng, et al.
Published: (2024)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
by: Xie, Jingchao, et al.
Published: (2025)
by: Xie, Jingchao, et al.
Published: (2025)
TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment
by: Qin, Chunxia, et al.
Published: (2026)
by: Qin, Chunxia, et al.
Published: (2026)
PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest
by: Beal, Josh, et al.
Published: (2026)
by: Beal, Josh, et al.
Published: (2026)
End-to-End Implicit Neural Representations for Classification
by: Gielisse, Alexander, et al.
Published: (2025)
by: Gielisse, Alexander, et al.
Published: (2025)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
End-to-End HOI Reconstruction Transformer with Graph-based Encoding
by: Wang, Zhenrong, et al.
Published: (2025)
by: Wang, Zhenrong, et al.
Published: (2025)
OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images
by: Zhao, Jiaqi, et al.
Published: (2024)
by: Zhao, Jiaqi, et al.
Published: (2024)
MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild
by: Fang, Xi, et al.
Published: (2024)
by: Fang, Xi, et al.
Published: (2024)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
OFMPNet: Deep End-to-End Model for Occupancy and Flow Prediction in Urban Environment
by: Murhij, Youshaa, et al.
Published: (2024)
by: Murhij, Youshaa, et al.
Published: (2024)
End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
by: Chu, Wenda, et al.
Published: (2026)
by: Chu, Wenda, et al.
Published: (2026)
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
by: Jiang, Yilei, et al.
Published: (2025)
by: Jiang, Yilei, et al.
Published: (2025)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Exploring the Causality of End-to-End Autonomous Driving
by: Li, Jiankun, et al.
Published: (2024)
by: Li, Jiankun, et al.
Published: (2024)
AD$^2$: Analysis and Detection of Adversarial Threats in Visual Perception for End-to-End Autonomous Driving Systems
by: Sahu, Ishan, et al.
Published: (2026)
by: Sahu, Ishan, et al.
Published: (2026)
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
by: Shao, Hao, et al.
Published: (2026)
by: Shao, Hao, et al.
Published: (2026)
OED: Towards One-stage End-to-End Dynamic Scene Graph Generation
by: Wang, Guan, et al.
Published: (2024)
by: Wang, Guan, et al.
Published: (2024)
Similar Items
-
PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
by: Mahadev, Rohan, et al.
Published: (2026) -
V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
by: Guhan, Pooja, et al.
Published: (2025) -
VSD-MOT: End-to-End Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Distillation
by: Du, Jun
Published: (2026) -
End-to-End Visual Autonomous Parking via Control-Aided Attention
by: Chen, Chao, et al.
Published: (2025) -
PAVE: An End-to-End Dataset for Production Autonomous Vehicle Evaluation
by: Li, Xiangyu, et al.
Published: (2025)