Vehicle-centric Perception via Multimodal Structured Pre-training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Wentao, Wang, Xiao, Li, Chenglong, Tang, Jin, Luo, Bin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Pre-training on High Definition X-ray Images: An Experimental Study
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
von: Wang, Xiao, et al.
Veröffentlicht: (2026)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
von: Tang, Yiwen, et al.
Veröffentlicht: (2023)
von: Tang, Yiwen, et al.
Veröffentlicht: (2023)
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models
von: Wu, Wentao, et al.
Veröffentlicht: (2024)
von: Wu, Wentao, et al.
Veröffentlicht: (2024)
Multi-modal Vision Pre-training for Medical Image Analysis
von: Rui, Shaohao, et al.
Veröffentlicht: (2024)
von: Rui, Shaohao, et al.
Veröffentlicht: (2024)
Stylized Structural Patterns for Improved Neural Network Pre-training
von: Salehi, Farnood, et al.
Veröffentlicht: (2025)
von: Salehi, Farnood, et al.
Veröffentlicht: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2024)
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2024)
Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute Recognition
von: Kong, Weizhe, et al.
Veröffentlicht: (2025)
von: Kong, Weizhe, et al.
Veröffentlicht: (2025)
Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
von: Zhu, Xun, et al.
Veröffentlicht: (2025)
von: Zhu, Xun, et al.
Veröffentlicht: (2025)
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
von: Ding, Renhua, et al.
Veröffentlicht: (2024)
von: Ding, Renhua, et al.
Veröffentlicht: (2024)
Gradient-based Fine-Tuning through Pre-trained Model Regularization
von: Liu, Xuanbo, et al.
Veröffentlicht: (2025)
von: Liu, Xuanbo, et al.
Veröffentlicht: (2025)
Dynamic Graph Induced Contour-aware Heat Conduction Network for Event-based Object Detection
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Dataset Ownership Verification in Contrastive Pre-trained Models
von: Xie, Yuechen, et al.
Veröffentlicht: (2025)
von: Xie, Yuechen, et al.
Veröffentlicht: (2025)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
von: Li, Liupeng, et al.
Veröffentlicht: (2026)
von: Li, Liupeng, et al.
Veröffentlicht: (2026)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
von: Lin, Ronghao, et al.
Veröffentlicht: (2024)
von: Lin, Ronghao, et al.
Veröffentlicht: (2024)
Self-supervised Pre-training of Text Recognizers
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
von: Chen, Hao, et al.
Veröffentlicht: (2023)
von: Chen, Hao, et al.
Veröffentlicht: (2023)
SugarcaneNet: An Optimized Ensemble of LASSO-Regularized Pre-trained Models for Accurate Disease Classification
von: Talukder, Md. Simul Hasan, et al.
Veröffentlicht: (2024)
von: Talukder, Md. Simul Hasan, et al.
Veröffentlicht: (2024)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
SlotPi: Physics-informed Object-centric Reasoning Models
von: Li, Jian, et al.
Veröffentlicht: (2025)
von: Li, Jian, et al.
Veröffentlicht: (2025)
Practical Continual Forgetting for Pre-trained Vision Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
Pre-training with Random Orthogonal Projection Image Modeling
von: Haghighat, Maryam, et al.
Veröffentlicht: (2023)
von: Haghighat, Maryam, et al.
Veröffentlicht: (2023)
Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
von: Reddy, N Dinesh, et al.
Veröffentlicht: (2025)
von: Reddy, N Dinesh, et al.
Veröffentlicht: (2025)
Enhanced Cooperative Perception for Autonomous Vehicles Using Imperfect Communication
von: Sarlak, Ahmad, et al.
Veröffentlicht: (2024)
von: Sarlak, Ahmad, et al.
Veröffentlicht: (2024)
Comparative Analysis of Deep Learning Models for Perception in Autonomous Vehicles
von: Khan, Jalal
Veröffentlicht: (2025)
von: Khan, Jalal
Veröffentlicht: (2025)
Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models
von: Peng, Bincheng, et al.
Veröffentlicht: (2026)
von: Peng, Bincheng, et al.
Veröffentlicht: (2026)
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
von: Zhang, Gengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Gengwei, et al.
Veröffentlicht: (2024)
Domain Guidance: A Simple Transfer Approach for a Pre-trained Diffusion Model
von: Zhong, Jincheng, et al.
Veröffentlicht: (2025)
von: Zhong, Jincheng, et al.
Veröffentlicht: (2025)
Pre-training Vision Transformers with Formula-driven Supervised Learning
von: Kataoka, Hirokatsu, et al.
Veröffentlicht: (2022)
von: Kataoka, Hirokatsu, et al.
Veröffentlicht: (2022)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
von: Berrada, Tariq, et al.
Veröffentlicht: (2023)
von: Berrada, Tariq, et al.
Veröffentlicht: (2023)
One Prompt Word is Enough to Boost Adversarial Robustness for Pre-trained Vision-Language Models
von: Li, Lin, et al.
Veröffentlicht: (2024)
von: Li, Lin, et al.
Veröffentlicht: (2024)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
von: Bawazir, Ameera, et al.
Veröffentlicht: (2024)
von: Bawazir, Ameera, et al.
Veröffentlicht: (2024)
Benchmarking the Influence of Pre-training on Explanation Performance in MR Image Classification
von: Oliveira, Marta, et al.
Veröffentlicht: (2023)
von: Oliveira, Marta, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
von: Wu, Wentao, et al.
Veröffentlicht: (2025) -
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
von: Wang, Xiao, et al.
Veröffentlicht: (2025) -
Pre-training on High Definition X-ray Images: An Experimental Study
von: Wang, Xiao, et al.
Veröffentlicht: (2024) -
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
von: Wang, Xiao, et al.
Veröffentlicht: (2026) -
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
von: Wu, Wentao, et al.
Veröffentlicht: (2025)