T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xiao, Wang, Ziwen, Kong, Weizhe, Wu, Wentao, Li, Yuehang, Zheng, Aihua, Li, Chenglong, Tang, Jin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Collaborative Enhancement Network for Low-quality Multi-spectral Vehicle Re-identification
by: Zheng, Aihua, et al.
Published: (2025)
by: Zheng, Aihua, et al.
Published: (2025)
State Space Model for New-Generation Network Alternative to Transformers: A Survey
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models
by: Wu, Wentao, et al.
Published: (2024)
by: Wu, Wentao, et al.
Published: (2024)
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
DCG ReID: Disentangling Collaboration and Guidance Fusion Representations for Multi-modal Vehicle Re-Identification
by: Zheng, Aihua, et al.
Published: (2026)
by: Zheng, Aihua, et al.
Published: (2026)
Pre-training on High Definition X-ray Images: An Experimental Study
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
An Empirical Study of Mamba-based Pedestrian Attribute Recognition
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Esports Debut as a Medal Event at 2023 Asian Games: Exploring Public Perceptions with BERTopic and GPT-4 Topic Fine-Tuning
by: Qian, Tyreal Yizhou, et al.
Published: (2024)
by: Qian, Tyreal Yizhou, et al.
Published: (2024)
Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute Recognition
by: Kong, Weizhe, et al.
Published: (2025)
by: Kong, Weizhe, et al.
Published: (2025)
SequencePAR: Understanding Pedestrian Attributes via A Sequence Generation Paradigm
by: Jin, Jiandong, et al.
Published: (2023)
by: Jin, Jiandong, et al.
Published: (2023)
Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies
by: Wang, Peiran, et al.
Published: (2024)
by: Wang, Peiran, et al.
Published: (2024)
ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
FITA: Fine-grained Image-Text Aligner for Radiology Report Generation
by: Yang, Honglong, et al.
Published: (2024)
by: Yang, Honglong, et al.
Published: (2024)
Fine-grained Image Retrieval via Dual-Vision Adaptation
by: Jiang, Xin, et al.
Published: (2025)
by: Jiang, Xin, et al.
Published: (2025)
Panoptic Perception: A Novel Task and Fine-grained Dataset for Universal Remote Sensing Image Interpretation
by: Zhao, Danpei, et al.
Published: (2024)
by: Zhao, Danpei, et al.
Published: (2024)
UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification
by: Wan, Xixi, et al.
Published: (2025)
by: Wan, Xixi, et al.
Published: (2025)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Decoupled Cross-Modal Alignment Network for Text-RGBT Person Retrieval and A High-Quality Benchmark
by: Deng, Yifei, et al.
Published: (2025)
by: Deng, Yifei, et al.
Published: (2025)
Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark
by: Deng, Yifei, et al.
Published: (2026)
by: Deng, Yifei, et al.
Published: (2026)
PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation
by: Yin, Yifan, et al.
Published: (2025)
by: Yin, Yifan, et al.
Published: (2025)
Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval
by: Wang, Shijie, et al.
Published: (2026)
by: Wang, Shijie, et al.
Published: (2026)
DehazeMamba: SAR-guided Optical Remote Sensing Image Dehazing with Adaptive State Space Model
by: Zhao, Zhicheng, et al.
Published: (2025)
by: Zhao, Zhicheng, et al.
Published: (2025)
Fine-grained Text to Image Synthesis
by: Ouyang, Xu, et al.
Published: (2024)
by: Ouyang, Xu, et al.
Published: (2024)
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering
by: Zhao, Zhicheng, et al.
Published: (2024)
by: Zhao, Zhicheng, et al.
Published: (2024)
Feynman Integral Reduction without Integration-By-Parts
by: Wang, Ziwen, et al.
Published: (2024)
by: Wang, Ziwen, et al.
Published: (2024)
Language-driven Fine-grained Retrieval
by: Wang, Shijie, et al.
Published: (2025)
by: Wang, Shijie, et al.
Published: (2025)
Bridge to Non-Barrier Communication: Gloss-Prompted Fine-grained Cued Speech Gesture Generation with Diffusion Model
by: Lei, Wentao, et al.
Published: (2024)
by: Lei, Wentao, et al.
Published: (2024)
Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images
by: Hu, Guyue, et al.
Published: (2026)
by: Hu, Guyue, et al.
Published: (2026)
DiRW: Path-Aware Digraph Learning for Heterophily
by: Su, Daohan, et al.
Published: (2024)
by: Su, Daohan, et al.
Published: (2024)
CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Parallel Augmentation and Dual Enhancement for Occluded Person Re-identification
by: Wang, Zi, et al.
Published: (2022)
by: Wang, Zi, et al.
Published: (2022)
Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
by: Lan, Bangxiang, et al.
Published: (2025)
by: Lan, Bangxiang, et al.
Published: (2025)
Sign Language Translation using Frame and Event Stream: Benchmark Dataset and Algorithms
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
by: He, Junwen, et al.
Published: (2024)
by: He, Junwen, et al.
Published: (2024)
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
by: Lin, Xingtao, et al.
Published: (2024)
by: Lin, Xingtao, et al.
Published: (2024)
Multi-label Text Classification using GloVe and Neural Network Models
by: Wang, Hongren
Published: (2023)
by: Wang, Hongren
Published: (2023)
Similar Items
-
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025) -
Collaborative Enhancement Network for Low-quality Multi-spectral Vehicle Re-identification
by: Zheng, Aihua, et al.
Published: (2025) -
State Space Model for New-Generation Network Alternative to Transformers: A Survey
by: Wang, Xiao, et al.
Published: (2024) -
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models
by: Wu, Wentao, et al.
Published: (2024) -
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
by: Wang, Xiao, et al.
Published: (2025)