HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Xinpeng, Han, Jianhua, Xu, Hang, Zhang, Wei, Li, Xiaomeng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
by: Ding, Xinpeng, et al.
Published: (2024)
by: Ding, Xinpeng, et al.
Published: (2024)
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
by: Huang, Runhui, et al.
Published: (2024)
by: Huang, Runhui, et al.
Published: (2024)
PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
by: Ding, Xinpeng, et al.
Published: (2025)
by: Ding, Xinpeng, et al.
Published: (2025)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
by: Liu, Xianjie, et al.
Published: (2025)
by: Liu, Xianjie, et al.
Published: (2025)
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
by: Nie, Ming, et al.
Published: (2023)
by: Nie, Ming, et al.
Published: (2023)
Hi-ResNet: Edge Detail Enhancement for High-Resolution Remote Sensing Segmentation
by: Chen, Yuxia, et al.
Published: (2023)
by: Chen, Yuxia, et al.
Published: (2023)
HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving
by: Lin, R. D., et al.
Published: (2025)
by: Lin, R. D., et al.
Published: (2025)
HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
by: Zhang, Ruifei, et al.
Published: (2025)
by: Zhang, Ruifei, et al.
Published: (2025)
HiP-AD: Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single Decoder
by: Tang, Yingqi, et al.
Published: (2025)
by: Tang, Yingqi, et al.
Published: (2025)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
by: Gao, Ruiyuan, et al.
Published: (2024)
by: Gao, Ruiyuan, et al.
Published: (2024)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
by: Han, Lawrence
Published: (2026)
by: Han, Lawrence
Published: (2026)
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
by: Luo, Bingjun, et al.
Published: (2026)
by: Luo, Bingjun, et al.
Published: (2026)
BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving
by: Chen, Zeming, et al.
Published: (2025)
by: Chen, Zeming, et al.
Published: (2025)
123D: Unifying Multi-Modal Autonomous Driving Data at Scale
by: Dauner, Daniel, et al.
Published: (2026)
by: Dauner, Daniel, et al.
Published: (2026)
HiLo: Detailed and Robust 3D Clothed Human Reconstruction with High-and Low-Frequency Information of Parametric Models
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROI
by: Reidy, Brendan, et al.
Published: (2024)
by: Reidy, Brendan, et al.
Published: (2024)
HiCMamba: Enhancing Hi-C Resolution and Identifying 3D Genome Structures with State Space Modeling
by: Yang, Minghao, et al.
Published: (2025)
by: Yang, Minghao, et al.
Published: (2025)
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
by: Nie, Ming, et al.
Published: (2026)
by: Nie, Ming, et al.
Published: (2026)
EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction
by: Cai, Han, et al.
Published: (2022)
by: Cai, Han, et al.
Published: (2022)
DriveLM: Driving with Graph Visual Question Answering
by: Sima, Chonghao, et al.
Published: (2023)
by: Sima, Chonghao, et al.
Published: (2023)
3D Surface Reconstruction with Enhanced High-Frequency Details
by: Zhang, Shikun, et al.
Published: (2025)
by: Zhang, Shikun, et al.
Published: (2025)
SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving
by: Li, Yiming, et al.
Published: (2023)
by: Li, Yiming, et al.
Published: (2023)
Detail-Enhancing Framework for Reference-Based Image Super-Resolution
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
by: Han, Jianhua, et al.
Published: (2025)
by: Han, Jianhua, et al.
Published: (2025)
HiLO: High-Level Object Fusion for Autonomous Driving using Transformers
by: Osterburg, Timo, et al.
Published: (2025)
by: Osterburg, Timo, et al.
Published: (2025)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
by: Lu, Junxin, et al.
Published: (2026)
by: Lu, Junxin, et al.
Published: (2026)
Precise Drive with VLM: First Prize Solution for PRCV 2024 Drive LM challenge
by: Huang, Bin, et al.
Published: (2024)
by: Huang, Bin, et al.
Published: (2024)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
Eliminating VAE for Fast and High-Resolution Generative Detail Restoration
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving
by: Chen, Runjian, et al.
Published: (2022)
by: Chen, Runjian, et al.
Published: (2022)
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
by: Zhou, Xirui, et al.
Published: (2025)
by: Zhou, Xirui, et al.
Published: (2025)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
by: Chen, Zisheng, et al.
Published: (2025)
by: Chen, Zisheng, et al.
Published: (2025)
Hi-Mamba: Hierarchical Mamba for Efficient Image Super-Resolution
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
HiMat: DiT-based Ultra-High Resolution SVBRDF Generation
by: Wang, Zixiong, et al.
Published: (2025)
by: Wang, Zixiong, et al.
Published: (2025)
HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
by: Liu, Yichen, et al.
Published: (2026)
by: Liu, Yichen, et al.
Published: (2026)
Similar Items
-
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
by: Ding, Xinpeng, et al.
Published: (2024) -
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
by: Huang, Runhui, et al.
Published: (2024) -
PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
by: Ding, Xinpeng, et al.
Published: (2025) -
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
by: Liu, Xianjie, et al.
Published: (2025) -
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
by: Nie, Ming, et al.
Published: (2023)