HENet: Hybrid Encoding for End-to-end Multi-task 3D Perception from Multi-view Cameras
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Zhongyu, Lin, ZhiWei, Wang, Xinhao, Wang, Yongtao, Xing, Yun, Qi, Shengxiang, Dong, Nan, Yang, Ming-Hsuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving
di: Xia, Zhongyu, et al.
Pubblicazione: (2025)
di: Xia, Zhongyu, et al.
Pubblicazione: (2025)
PTQAT: A Hybrid Parameter-Efficient Quantization Algorithm for 3D Perception Tasks
di: Wang, Xinhao, et al.
Pubblicazione: (2025)
di: Wang, Xinhao, et al.
Pubblicazione: (2025)
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
di: Lin, Zhiwei, et al.
Pubblicazione: (2022)
di: Lin, Zhiwei, et al.
Pubblicazione: (2022)
RCBEVDet: Radar-camera Fusion in Bird's Eye View for 3D Object Detection
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
OpenAD: Open-World Autonomous Driving Benchmark for 3D Object Detection
di: Xia, Zhongyu, et al.
Pubblicazione: (2024)
di: Xia, Zhongyu, et al.
Pubblicazione: (2024)
InsFusion: Rethink Instance-level LiDAR-Camera Fusion for 3D Object Detection
di: Xia, Zhongyu, et al.
Pubblicazione: (2025)
di: Xia, Zhongyu, et al.
Pubblicazione: (2025)
KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
di: Xia, Zhongyu, et al.
Pubblicazione: (2025)
di: Xia, Zhongyu, et al.
Pubblicazione: (2025)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
di: Yang, Jiawei, et al.
Pubblicazione: (2025)
di: Yang, Jiawei, et al.
Pubblicazione: (2025)
AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation
di: Savigny, Henri, et al.
Pubblicazione: (2025)
di: Savigny, Henri, et al.
Pubblicazione: (2025)
HeLoFusion: An Efficient and Scalable Encoder for Modeling Heterogeneous and Multi-Scale Interactions in Trajectory Prediction
di: Wei, Bingqing, et al.
Pubblicazione: (2025)
di: Wei, Bingqing, et al.
Pubblicazione: (2025)
R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance
di: Guo, Ziang, et al.
Pubblicazione: (2024)
di: Guo, Ziang, et al.
Pubblicazione: (2024)
MEBS: Multi-task End-to-end Bid Shading for Multi-slot Display Advertising
di: Gong, Zhen, et al.
Pubblicazione: (2024)
di: Gong, Zhen, et al.
Pubblicazione: (2024)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
di: Zhou, Xiaoyu, et al.
Pubblicazione: (2025)
di: Zhou, Xiaoyu, et al.
Pubblicazione: (2025)
Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
di: Ivanovic, Boris, et al.
Pubblicazione: (2025)
di: Ivanovic, Boris, et al.
Pubblicazione: (2025)
Efficiently Disentangling CLIP for Multi-Object Perception
di: Rawlekar, Samyak, et al.
Pubblicazione: (2025)
di: Rawlekar, Samyak, et al.
Pubblicazione: (2025)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
di: Wang, Xinhao, et al.
Pubblicazione: (2026)
di: Wang, Xinhao, et al.
Pubblicazione: (2026)
TEOcc: Radar-camera Multi-modal Occupancy Prediction via Temporal Enhancement
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
Auto-Configured Networks for Multi-Scale Multi-Output Time-Series Forecasting
di: Zha, Yumeng, et al.
Pubblicazione: (2026)
di: Zha, Yumeng, et al.
Pubblicazione: (2026)
RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
di: Yue, Jingtong, et al.
Pubblicazione: (2025)
di: Yue, Jingtong, et al.
Pubblicazione: (2025)
COMICS: End-to-end Bi-grained Contrastive Learning for Multi-face Forgery Detection
di: Zhang, Cong, et al.
Pubblicazione: (2023)
di: Zhang, Cong, et al.
Pubblicazione: (2023)
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
di: Zhao, Xiran, et al.
Pubblicazione: (2026)
di: Zhao, Xiran, et al.
Pubblicazione: (2026)
Multi-task Image Restoration Guided By Robust DINO Features
di: Lin, Xin, et al.
Pubblicazione: (2023)
di: Lin, Xin, et al.
Pubblicazione: (2023)
End-to-end Autonomous Vehicle Following System using Monocular Fisheye Camera
di: Zhang, Jiale, et al.
Pubblicazione: (2025)
di: Zhang, Jiale, et al.
Pubblicazione: (2025)
HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
RCBEVDet++: Toward High-accuracy Radar-Camera Fusion 3D Perception Network
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
di: Wang, Yukun, et al.
Pubblicazione: (2026)
di: Wang, Yukun, et al.
Pubblicazione: (2026)
FedHENet: A Frugal Federated Learning Framework for Heterogeneous Environments
di: Dopico-Castro, Alejandro, et al.
Pubblicazione: (2026)
di: Dopico-Castro, Alejandro, et al.
Pubblicazione: (2026)
Multi-task Learning for Heterogeneous Data via Integrating Shared and Task-Specific Encodings
di: Sui, Yang, et al.
Pubblicazione: (2025)
di: Sui, Yang, et al.
Pubblicazione: (2025)
DV-3DLane: End-to-end Multi-modal 3D Lane Detection with Dual-view Representation
di: Luo, Yueru, et al.
Pubblicazione: (2024)
di: Luo, Yueru, et al.
Pubblicazione: (2024)
MDDM: A Multi-view Discriminative Enhanced Diffusion-based Model for Speech Enhancement
di: Xu, Nan, et al.
Pubblicazione: (2025)
di: Xu, Nan, et al.
Pubblicazione: (2025)
Multi-view Disentanglement for Reinforcement Learning with Multiple Cameras
di: Dunion, Mhairi, et al.
Pubblicazione: (2024)
di: Dunion, Mhairi, et al.
Pubblicazione: (2024)
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding
di: Wang, Jieyi, et al.
Pubblicazione: (2026)
di: Wang, Jieyi, et al.
Pubblicazione: (2026)
Multi-level Reliable Guidance for Unpaired Multi-view Clustering
di: Xin, Like, et al.
Pubblicazione: (2024)
di: Xin, Like, et al.
Pubblicazione: (2024)
SEMPose: A Single End-to-end Network for Multi-object Pose Estimation
di: Liu, Xin, et al.
Pubblicazione: (2024)
di: Liu, Xin, et al.
Pubblicazione: (2024)
ParkingE2E: Camera-based End-to-end Parking Network, from Images to Planning
di: Li, Changze, et al.
Pubblicazione: (2024)
di: Li, Changze, et al.
Pubblicazione: (2024)
Which Side Are You On? A Multi-task Dataset for End-to-End Argument Summarisation and Evaluation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
di: Zheng, Chaoda, et al.
Pubblicazione: (2026)
di: Zheng, Chaoda, et al.
Pubblicazione: (2026)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention
di: Xu, Dejia, et al.
Pubblicazione: (2024)
di: Xu, Dejia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving
di: Xia, Zhongyu, et al.
Pubblicazione: (2025) -
PTQAT: A Hybrid Parameter-Efficient Quantization Algorithm for 3D Perception Tasks
di: Wang, Xinhao, et al.
Pubblicazione: (2025) -
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
di: Lin, Zhiwei, et al.
Pubblicazione: (2022) -
RCBEVDet: Radar-camera Fusion in Bird's Eye View for 3D Object Detection
di: Lin, Zhiwei, et al.
Pubblicazione: (2024) -
OpenAD: Open-World Autonomous Driving Benchmark for 3D Object Detection
di: Xia, Zhongyu, et al.
Pubblicazione: (2024)