CLIP-BEVFormer: Enhancing Multi-View Image-Based BEV Detector with Ground Truth Flow
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Chenbin, Yaman, Burhaneddin, Velipasalar, Senem, Ren, Liu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VLP: Vision Language Planning for Autonomous Driving
von: Pan, Chenbin, et al.
Veröffentlicht: (2024)
von: Pan, Chenbin, et al.
Veröffentlicht: (2024)
BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
von: Ye, Xin, et al.
Veröffentlicht: (2025)
von: Ye, Xin, et al.
Veröffentlicht: (2025)
DG-MVP: 3D Domain Generalization via Multiple Views of Point Clouds for Classification
von: Ren, Huantao, et al.
Veröffentlicht: (2025)
von: Ren, Huantao, et al.
Veröffentlicht: (2025)
GaitPoint+: A Gait Recognition Network Incorporating Point Cloud Analysis and Recycling
von: Ren, Huantao, et al.
Veröffentlicht: (2024)
von: Ren, Huantao, et al.
Veröffentlicht: (2024)
3D-PointZshotS: Geometry-Aware 3D Point Cloud Zero-Shot Semantic Segmentation Narrowing the Visual-Semantic Gap
von: Yang, Minmin, et al.
Veröffentlicht: (2025)
von: Yang, Minmin, et al.
Veröffentlicht: (2025)
Trans${^2}$-CBCT: A Dual-Transformer Framework for Sparse-View CBCT Reconstruction
von: Yang, Minmin, et al.
Veröffentlicht: (2025)
von: Yang, Minmin, et al.
Veröffentlicht: (2025)
Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting
von: Lu, Yiren, et al.
Veröffentlicht: (2026)
von: Lu, Yiren, et al.
Veröffentlicht: (2026)
PRISM: Product Retrieval In Shopping Carts using Hybrid Matching
von: Kabadayi, Arda, et al.
Veröffentlicht: (2025)
von: Kabadayi, Arda, et al.
Veröffentlicht: (2025)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
von: Ma, Yunsheng, et al.
Veröffentlicht: (2024)
von: Ma, Yunsheng, et al.
Veröffentlicht: (2024)
LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool
von: Ma, Yue, et al.
Veröffentlicht: (2024)
von: Ma, Yue, et al.
Veröffentlicht: (2024)
PaPr: Training-Free One-Step Patch Pruning with Lightweight ConvNets for Faster Inference
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
Only My Model On My Data: A Privacy Preserving Approach Protecting one Model and Deceiving Unauthorized Black-Box Models
von: Chai, Weiheng, et al.
Veröffentlicht: (2024)
von: Chai, Weiheng, et al.
Veröffentlicht: (2024)
OccTransformer: Improving BEVFormer for 3D camera-only occupancy prediction
von: Liu, Jian, et al.
Veröffentlicht: (2024)
von: Liu, Jian, et al.
Veröffentlicht: (2024)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
Beyond the Ground Truth: Enhanced Supervision for Image Restoration
von: Ryou, Donghun, et al.
Veröffentlicht: (2025)
von: Ryou, Donghun, et al.
Veröffentlicht: (2025)
Leveraging BEV Paradigm for Ground-to-Aerial Image Synthesis
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
Calibrated and Resource-Aware Super-Resolution for Reliable Driver Behavior Analysis
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
von: Ma, Yunsheng, et al.
Veröffentlicht: (2025)
von: Ma, Yunsheng, et al.
Veröffentlicht: (2025)
SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
HV-BEV: Decoupling Horizontal and Vertical Feature Sampling for Multi-View 3D Object Detection
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping
von: Wei, Jiale, et al.
Veröffentlicht: (2024)
von: Wei, Jiale, et al.
Veröffentlicht: (2024)
Benchmarking Multi-View BEV Object Detection with Mixed Pinhole and Fisheye Cameras
von: Liu, Xiangzhong, et al.
Veröffentlicht: (2026)
von: Liu, Xiangzhong, et al.
Veröffentlicht: (2026)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
von: Song, Dan, et al.
Veröffentlicht: (2023)
von: Song, Dan, et al.
Veröffentlicht: (2023)
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
von: Pan, Chenbin, et al.
Veröffentlicht: (2025)
von: Pan, Chenbin, et al.
Veröffentlicht: (2025)
Duoduo CLIP: Efficient 3D Understanding with Multi-View Images
von: Lee, Han-Hung, et al.
Veröffentlicht: (2024)
von: Lee, Han-Hung, et al.
Veröffentlicht: (2024)
Minimizing Occlusion Effect on Multi-View Camera Perception in BEV with Multi-Sensor Fusion
von: Kumar, Sanjay, et al.
Veröffentlicht: (2025)
von: Kumar, Sanjay, et al.
Veröffentlicht: (2025)
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy
von: Zhang, Zaibin, et al.
Veröffentlicht: (2023)
von: Zhang, Zaibin, et al.
Veröffentlicht: (2023)
Sparse BEV Fusion with Self-View Consistency for Multi-View Detection and Tracking
von: Toida, Keisuke, et al.
Veröffentlicht: (2025)
von: Toida, Keisuke, et al.
Veröffentlicht: (2025)
FSF-Net: Enhance 4D Occupancy Forecasting with Coarse BEV Scene Flow for Autonomous Driving
von: Guo, Erxin, et al.
Veröffentlicht: (2024)
von: Guo, Erxin, et al.
Veröffentlicht: (2024)
RopeBEV: A Multi-Camera Roadside Perception Network in Bird's-Eye-View
von: Jia, Jinrang, et al.
Veröffentlicht: (2024)
von: Jia, Jinrang, et al.
Veröffentlicht: (2024)
CLIP3D-AD: Extending CLIP for 3D Few-Shot Anomaly Detection with Multi-View Images Generation
von: Zuo, Zuo, et al.
Veröffentlicht: (2024)
von: Zuo, Zuo, et al.
Veröffentlicht: (2024)
BEV$^2$PR: BEV-Enhanced Visual Place Recognition with Structural Cues
von: Ge, Fudong, et al.
Veröffentlicht: (2024)
von: Ge, Fudong, et al.
Veröffentlicht: (2024)
GeoBEV: Learning Geometric BEV Representation for Multi-view 3D Object Detection
von: Zhang, Jinqing, et al.
Veröffentlicht: (2024)
von: Zhang, Jinqing, et al.
Veröffentlicht: (2024)
Adapting Fine-Grained Cross-View Localization to Areas without Fine Ground Truth
von: Xia, Zimin, et al.
Veröffentlicht: (2024)
von: Xia, Zimin, et al.
Veröffentlicht: (2024)
BrainMCLIP: Brain Image Decoding with Multi-Layer feature Fusion of CLIP
von: Xia, Tian, et al.
Veröffentlicht: (2025)
von: Xia, Tian, et al.
Veröffentlicht: (2025)
SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors
von: Li, Aixuan, et al.
Veröffentlicht: (2025)
von: Li, Aixuan, et al.
Veröffentlicht: (2025)
RoadBEV: Road Surface Reconstruction in Bird's Eye View
von: Zhao, Tong, et al.
Veröffentlicht: (2024)
von: Zhao, Tong, et al.
Veröffentlicht: (2024)
DualBEV: Unifying Dual View Transformation with Probabilistic Correspondences
von: Li, Peidong, et al.
Veröffentlicht: (2024)
von: Li, Peidong, et al.
Veröffentlicht: (2024)
Not Just Streaks: Towards Ground Truth for Single Image Deraining
von: Ba, Yunhao, et al.
Veröffentlicht: (2022)
von: Ba, Yunhao, et al.
Veröffentlicht: (2022)
GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
von: Song, Ziying, et al.
Veröffentlicht: (2024)
von: Song, Ziying, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VLP: Vision Language Planning for Autonomous Driving
von: Pan, Chenbin, et al.
Veröffentlicht: (2024) -
BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
von: Ye, Xin, et al.
Veröffentlicht: (2025) -
DG-MVP: 3D Domain Generalization via Multiple Views of Point Clouds for Classification
von: Ren, Huantao, et al.
Veröffentlicht: (2025) -
GaitPoint+: A Gait Recognition Network Incorporating Point Cloud Analysis and Recycling
von: Ren, Huantao, et al.
Veröffentlicht: (2024) -
3D-PointZshotS: Geometry-Aware 3D Point Cloud Zero-Shot Semantic Segmentation Narrowing the Visual-Semantic Gap
von: Yang, Minmin, et al.
Veröffentlicht: (2025)