Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
Fuente:
arXiv
Salvato in:
| Autori principali: | Fu, Luxuan, Liu, Chong, Yang, Bisheng, Dong, Zhen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SVII-3D: Advancing Roadside Infrastructure Inventory with Decimeter-level 3D Localization and Comprehension from Sparse Street Imagery
di: Liu, Chong, et al.
Pubblicazione: (2026)
di: Liu, Chong, et al.
Pubblicazione: (2026)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
di: Jiao, Yang, et al.
Pubblicazione: (2024)
di: Jiao, Yang, et al.
Pubblicazione: (2024)
2.5D Object Detection for Intelligent Roadside Infrastructure
di: Polley, Nikolai, et al.
Pubblicazione: (2025)
di: Polley, Nikolai, et al.
Pubblicazione: (2025)
ME-CPT: Multi-Task Enhanced Cross-Temporal Point Transformer for Urban 3D Change Detection
di: Zhang, Luqi, et al.
Pubblicazione: (2025)
di: Zhang, Luqi, et al.
Pubblicazione: (2025)
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
di: Fang, Rongyao, et al.
Pubblicazione: (2025)
di: Fang, Rongyao, et al.
Pubblicazione: (2025)
GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting
di: Peng, Yuning, et al.
Pubblicazione: (2024)
di: Peng, Yuning, et al.
Pubblicazione: (2024)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
DPG-CD: Depth-Prior-Guided Cross-Modal Joint 2D-3D Change Detection
di: Zhang, Luqi, et al.
Pubblicazione: (2026)
di: Zhang, Luqi, et al.
Pubblicazione: (2026)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
di: Chen, Jiabin, et al.
Pubblicazione: (2025)
di: Chen, Jiabin, et al.
Pubblicazione: (2025)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
di: Li, Hengzhuang, et al.
Pubblicazione: (2025)
di: Li, Hengzhuang, et al.
Pubblicazione: (2025)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
di: Guan, Runwei, et al.
Pubblicazione: (2025)
di: Guan, Runwei, et al.
Pubblicazione: (2025)
Multimodal HD Mapping for Intersections by Intelligent Roadside Units
di: Chen, Zhongzhang, et al.
Pubblicazione: (2025)
di: Chen, Zhongzhang, et al.
Pubblicazione: (2025)
SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
di: Yue, Tongtian, et al.
Pubblicazione: (2024)
di: Yue, Tongtian, et al.
Pubblicazione: (2024)
RoScenes: A Large-scale Multi-view 3D Dataset for Roadside Perception
di: Zhu, Xiaosu, et al.
Pubblicazione: (2024)
di: Zhu, Xiaosu, et al.
Pubblicazione: (2024)
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
di: Fu, Teng, et al.
Pubblicazione: (2025)
di: Fu, Teng, et al.
Pubblicazione: (2025)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
di: Liu, Yifan, et al.
Pubblicazione: (2025)
di: Liu, Yifan, et al.
Pubblicazione: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
di: Lin, Junyan, et al.
Pubblicazione: (2026)
di: Lin, Junyan, et al.
Pubblicazione: (2026)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
VistaDream: Sampling multiview consistent images for single-view scene reconstruction
di: Wang, Haiping, et al.
Pubblicazione: (2024)
di: Wang, Haiping, et al.
Pubblicazione: (2024)
MoniRefer: A Real-world Large-scale Multi-modal Dataset based on Roadside Infrastructure for 3D Visual Grounding
di: Yang, Panquan, et al.
Pubblicazione: (2025)
di: Yang, Panquan, et al.
Pubblicazione: (2025)
RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground Simulation
di: Du, Yuwen, et al.
Pubblicazione: (2025)
di: Du, Yuwen, et al.
Pubblicazione: (2025)
FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators
di: Wang, Haiping, et al.
Pubblicazione: (2023)
di: Wang, Haiping, et al.
Pubblicazione: (2023)
Explicitly Guided Information Interaction Network for Cross-modal Point Cloud Completion
di: Xu, Hang, et al.
Pubblicazione: (2024)
di: Xu, Hang, et al.
Pubblicazione: (2024)
Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
di: Meng, Siyuan, et al.
Pubblicazione: (2026)
di: Meng, Siyuan, et al.
Pubblicazione: (2026)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
di: Choi, In Chong, et al.
Pubblicazione: (2026)
di: Choi, In Chong, et al.
Pubblicazione: (2026)
CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
di: An, Xiao, et al.
Pubblicazione: (2024)
di: An, Xiao, et al.
Pubblicazione: (2024)
Evaluating Graphical Perception Capabilities of Vision Transformers
di: Poonam, Poonam, et al.
Pubblicazione: (2026)
di: Poonam, Poonam, et al.
Pubblicazione: (2026)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
di: Wang, Shijing, et al.
Pubblicazione: (2025)
di: Wang, Shijing, et al.
Pubblicazione: (2025)
Accurate Cooperative Localization Utilizing LiDAR-equipped Roadside Infrastructure for Autonomous Driving
di: Jiang, Yuze, et al.
Pubblicazione: (2024)
di: Jiang, Yuze, et al.
Pubblicazione: (2024)
CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks
di: Wang, Beibei, et al.
Pubblicazione: (2024)
di: Wang, Beibei, et al.
Pubblicazione: (2024)
Unleashing Vision-Language Semantics for Deepfake Video Detection
di: Zhu, Jiawen, et al.
Pubblicazione: (2026)
di: Zhu, Jiawen, et al.
Pubblicazione: (2026)
Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
di: Yu, Peipeng, et al.
Pubblicazione: (2025)
di: Yu, Peipeng, et al.
Pubblicazione: (2025)
Expert Knowledge-Guided Decision Calibration for Accurate Fine-Grained Tree Species Classification
di: Long, Chen, et al.
Pubblicazione: (2026)
di: Long, Chen, et al.
Pubblicazione: (2026)
LifelongPR: Lifelong point cloud place recognition based on sample replay and prompt learning
di: Zou, Xianghong, et al.
Pubblicazione: (2025)
di: Zou, Xianghong, et al.
Pubblicazione: (2025)
RopeBEV: A Multi-Camera Roadside Perception Network in Bird's-Eye-View
di: Jia, Jinrang, et al.
Pubblicazione: (2024)
di: Jia, Jinrang, et al.
Pubblicazione: (2024)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
di: Tang, Liyan, et al.
Pubblicazione: (2025)
di: Tang, Liyan, et al.
Pubblicazione: (2025)
SGV3D:Towards Scenario Generalization for Vision-based Roadside 3D Object Detection
di: Yang, Lei, et al.
Pubblicazione: (2024)
di: Yang, Lei, et al.
Pubblicazione: (2024)
Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models
di: Hu, Rui, et al.
Pubblicazione: (2025)
di: Hu, Rui, et al.
Pubblicazione: (2025)
DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
di: Hu, Ying, et al.
Pubblicazione: (2024)
di: Hu, Ying, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SVII-3D: Advancing Roadside Infrastructure Inventory with Decimeter-level 3D Localization and Comprehension from Sparse Street Imagery
di: Liu, Chong, et al.
Pubblicazione: (2026) -
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
di: Jiao, Yang, et al.
Pubblicazione: (2024) -
2.5D Object Detection for Intelligent Roadside Infrastructure
di: Polley, Nikolai, et al.
Pubblicazione: (2025) -
ME-CPT: Multi-Task Enhanced Cross-Temporal Point Transformer for Urban 3D Change Detection
di: Zhang, Luqi, et al.
Pubblicazione: (2025) -
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
di: Fang, Rongyao, et al.
Pubblicazione: (2025)