Vehicle-centric Perception via Multimodal Structured Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Wentao, Wang, Xiao, Li, Chenglong, Tang, Jin, Luo, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Pre-training on High Definition X-ray Images: An Experimental Study
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
by: Tang, Yiwen, et al.
Published: (2023)
by: Tang, Yiwen, et al.
Published: (2023)
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models
by: Wu, Wentao, et al.
Published: (2024)
by: Wu, Wentao, et al.
Published: (2024)
Multi-modal Vision Pre-training for Medical Image Analysis
by: Rui, Shaohao, et al.
Published: (2024)
by: Rui, Shaohao, et al.
Published: (2024)
Stylized Structural Patterns for Improved Neural Network Pre-training
by: Salehi, Farnood, et al.
Published: (2025)
by: Salehi, Farnood, et al.
Published: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
by: Lai, Zhengfeng, et al.
Published: (2024)
by: Lai, Zhengfeng, et al.
Published: (2024)
Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute Recognition
by: Kong, Weizhe, et al.
Published: (2025)
by: Kong, Weizhe, et al.
Published: (2025)
Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
by: Zhu, Xun, et al.
Published: (2025)
by: Zhu, Xun, et al.
Published: (2025)
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
by: Ding, Renhua, et al.
Published: (2024)
by: Ding, Renhua, et al.
Published: (2024)
Gradient-based Fine-Tuning through Pre-trained Model Regularization
by: Liu, Xuanbo, et al.
Published: (2025)
by: Liu, Xuanbo, et al.
Published: (2025)
Dynamic Graph Induced Contour-aware Heat Conduction Network for Event-based Object Detection
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Dataset Ownership Verification in Contrastive Pre-trained Models
by: Xie, Yuechen, et al.
Published: (2025)
by: Xie, Yuechen, et al.
Published: (2025)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
by: Li, Liupeng, et al.
Published: (2026)
by: Li, Liupeng, et al.
Published: (2026)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
Self-supervised Pre-training of Text Recognizers
by: Kišš, Martin, et al.
Published: (2024)
by: Kišš, Martin, et al.
Published: (2024)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
by: Chen, Hao, et al.
Published: (2023)
by: Chen, Hao, et al.
Published: (2023)
SugarcaneNet: An Optimized Ensemble of LASSO-Regularized Pre-trained Models for Accurate Disease Classification
by: Talukder, Md. Simul Hasan, et al.
Published: (2024)
by: Talukder, Md. Simul Hasan, et al.
Published: (2024)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
SlotPi: Physics-informed Object-centric Reasoning Models
by: Li, Jian, et al.
Published: (2025)
by: Li, Jian, et al.
Published: (2025)
Practical Continual Forgetting for Pre-trained Vision Models
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Pre-training with Random Orthogonal Projection Image Modeling
by: Haghighat, Maryam, et al.
Published: (2023)
by: Haghighat, Maryam, et al.
Published: (2023)
Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
by: Reddy, N Dinesh, et al.
Published: (2025)
by: Reddy, N Dinesh, et al.
Published: (2025)
Enhanced Cooperative Perception for Autonomous Vehicles Using Imperfect Communication
by: Sarlak, Ahmad, et al.
Published: (2024)
by: Sarlak, Ahmad, et al.
Published: (2024)
Comparative Analysis of Deep Learning Models for Perception in Autonomous Vehicles
by: Khan, Jalal
Published: (2025)
by: Khan, Jalal
Published: (2025)
Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models
by: Peng, Bincheng, et al.
Published: (2026)
by: Peng, Bincheng, et al.
Published: (2026)
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
by: Zhang, Gengwei, et al.
Published: (2024)
by: Zhang, Gengwei, et al.
Published: (2024)
Domain Guidance: A Simple Transfer Approach for a Pre-trained Diffusion Model
by: Zhong, Jincheng, et al.
Published: (2025)
by: Zhong, Jincheng, et al.
Published: (2025)
Pre-training Vision Transformers with Formula-driven Supervised Learning
by: Kataoka, Hirokatsu, et al.
Published: (2022)
by: Kataoka, Hirokatsu, et al.
Published: (2022)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
by: Berrada, Tariq, et al.
Published: (2023)
by: Berrada, Tariq, et al.
Published: (2023)
One Prompt Word is Enough to Boost Adversarial Robustness for Pre-trained Vision-Language Models
by: Li, Lin, et al.
Published: (2024)
by: Li, Lin, et al.
Published: (2024)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
by: Bawazir, Ameera, et al.
Published: (2024)
by: Bawazir, Ameera, et al.
Published: (2024)
Benchmarking the Influence of Pre-training on Explanation Performance in MR Image Classification
by: Oliveira, Marta, et al.
Published: (2023)
by: Oliveira, Marta, et al.
Published: (2023)
Similar Items
-
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
by: Wu, Wentao, et al.
Published: (2025) -
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
by: Wang, Xiao, et al.
Published: (2025) -
Pre-training on High Definition X-ray Images: An Experimental Study
by: Wang, Xiao, et al.
Published: (2024) -
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
by: Wang, Xiao, et al.
Published: (2026) -
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
by: Wu, Wentao, et al.
Published: (2025)