Referring Expression Instance Retrieval and A Strong End-to-End Baseline
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Xiangzhao, Zhu, Kuan, Guo, Hongyu, Guo, Haiyun, Jiang, Ning, Lu, Quan, Tang, Ming, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
by: Guo, Hongyu, et al.
Published: (2025)
by: Guo, Hongyu, et al.
Published: (2025)
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
by: Hao, Xiangzhao, et al.
Published: (2026)
by: Hao, Xiangzhao, et al.
Published: (2026)
WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
by: Wang, Tianyue, et al.
Published: (2026)
by: Wang, Tianyue, et al.
Published: (2026)
ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
by: Yang, Tianyu, et al.
Published: (2026)
by: Yang, Tianyu, et al.
Published: (2026)
FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
by: An, Hongyan, et al.
Published: (2025)
by: An, Hongyan, et al.
Published: (2025)
AAformer: Auto-Aligned Transformer for Person Re-Identification
by: Zhu, Kuan, et al.
Published: (2021)
by: Zhu, Kuan, et al.
Published: (2021)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
by: He, Chenwei, et al.
Published: (2026)
by: He, Chenwei, et al.
Published: (2026)
Monocular Lane Detection Based on Deep Learning: A Survey
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
by: Wu, Changli, et al.
Published: (2024)
by: Wu, Changli, et al.
Published: (2024)
Reference Twice: A Simple and Unified Baseline for Few-Shot Instance Segmentation
by: Han, Yue, et al.
Published: (2023)
by: Han, Yue, et al.
Published: (2023)
End-to-End Human Instance Matting
by: Liu, Qinglin, et al.
Published: (2024)
by: Liu, Qinglin, et al.
Published: (2024)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
by: He, Jinghan, et al.
Published: (2024)
by: He, Jinghan, et al.
Published: (2024)
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
by: Tang, Wenhao, et al.
Published: (2025)
by: Tang, Wenhao, et al.
Published: (2025)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
STORM: End-to-End Referring Multi-Object Tracking in Videos
by: Lu, Zijia, et al.
Published: (2026)
by: Lu, Zijia, et al.
Published: (2026)
S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object Tracking
by: Tang, Tao, et al.
Published: (2024)
by: Tang, Tao, et al.
Published: (2024)
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation
by: Chen, Wenxin, et al.
Published: (2025)
by: Chen, Wenxin, et al.
Published: (2025)
InterMesh: Explicit Interaction-Aware End-to-End Multi-Person Human Mesh Recovery
by: Zheng, Kaili, et al.
Published: (2026)
by: Zheng, Kaili, et al.
Published: (2026)
ContourFormer: Real-Time Contour-Based End-to-End Instance Segmentation Transformer
by: Yao, Weiwei, et al.
Published: (2025)
by: Yao, Weiwei, et al.
Published: (2025)
UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery
by: Zhang, Huaxiang, et al.
Published: (2025)
by: Zhang, Huaxiang, et al.
Published: (2025)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
by: Han, Jianhua, et al.
Published: (2025)
by: Han, Jianhua, et al.
Published: (2025)
MaskFuser: Masked Fusion of Joint Multi-Modal Tokenization for End-to-End Autonomous Driving
by: Duan, Yiqun, et al.
Published: (2024)
by: Duan, Yiqun, et al.
Published: (2024)
Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning
by: Song, Yuehao, et al.
Published: (2026)
by: Song, Yuehao, et al.
Published: (2026)
OpenStereo: A Comprehensive Benchmark for Stereo Matching and Strong Baseline
by: Guo, Xianda, et al.
Published: (2023)
by: Guo, Xianda, et al.
Published: (2023)
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
by: Hao, Xiangzhao, et al.
Published: (2026)
by: Hao, Xiangzhao, et al.
Published: (2026)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
GenAD: Generative End-to-End Autonomous Driving
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving
by: Song, Ziying, et al.
Published: (2025)
by: Song, Ziying, et al.
Published: (2025)
Text-to-CAD Retrieval: a Strong Baseline
by: Pan, Honghu, et al.
Published: (2026)
by: Pan, Honghu, et al.
Published: (2026)
DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
Towards Fully Decoupled End-to-End Person Search
by: Zhang, Pengcheng, et al.
Published: (2023)
by: Zhang, Pengcheng, et al.
Published: (2023)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
End-to-End Autonomous Driving without Costly Modularization and 3D Manual Annotation
by: Guo, Mingzhe, et al.
Published: (2024)
by: Guo, Mingzhe, et al.
Published: (2024)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
SpaRC-AD: A Baseline for Radar-Camera Fusion in End-to-End Autonomous Driving
by: Wolters, Philipp, et al.
Published: (2025)
by: Wolters, Philipp, et al.
Published: (2025)
Differentiable NMS via Sinkhorn Matching for End-to-End Fabric Defect Detection
by: Lu, Zhengyang, et al.
Published: (2025)
by: Lu, Zhengyang, et al.
Published: (2025)
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
by: Zhu, Feng, et al.
Published: (2026)
by: Zhu, Feng, et al.
Published: (2026)
UniUncer: Unified Dynamic Static Uncertainty for End to End Driving
by: Gao, Yu, et al.
Published: (2026)
by: Gao, Yu, et al.
Published: (2026)
BEEP3D: Box-Supervised End-to-End Pseudo-Mask Generation for 3D Instance Segmentation
by: Yoo, Youngju, et al.
Published: (2025)
by: Yoo, Youngju, et al.
Published: (2025)
Similar Items
-
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
by: Guo, Hongyu, et al.
Published: (2025) -
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
by: Hao, Xiangzhao, et al.
Published: (2026) -
WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
by: Wang, Tianyue, et al.
Published: (2026) -
ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
by: Yang, Tianyu, et al.
Published: (2026) -
FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
by: An, Hongyan, et al.
Published: (2025)