3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Yiping, Li, Jinpeng, Ke, Wenyu, Luo, Yang, Ouyang, Jie, He, Zhongjie, Liu, Li, Fan, Hongchao, Wu, Hao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
di: Huang, Keli, et al.
Pubblicazione: (2022)
di: Huang, Keli, et al.
Pubblicazione: (2022)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
di: Han, Yudong, et al.
Pubblicazione: (2024)
di: Han, Yudong, et al.
Pubblicazione: (2024)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
di: Mehta, Vinit, et al.
Pubblicazione: (2025)
di: Mehta, Vinit, et al.
Pubblicazione: (2025)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
di: Li, Huibin, et al.
Pubblicazione: (2025)
di: Li, Huibin, et al.
Pubblicazione: (2025)
Collaborative AI Enhances Image Understanding in Materials Science
di: Yin, Ruoyan Avery, et al.
Pubblicazione: (2025)
di: Yin, Ruoyan Avery, et al.
Pubblicazione: (2025)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
di: Tran, Duc Dang Trung, et al.
Pubblicazione: (2024)
di: Tran, Duc Dang Trung, et al.
Pubblicazione: (2024)
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
di: Chen, Junting, et al.
Pubblicazione: (2025)
di: Chen, Junting, et al.
Pubblicazione: (2025)
Context-Aware Indoor Point Cloud Object Generation through User Instructions
di: Luo, Yiyang, et al.
Pubblicazione: (2023)
di: Luo, Yiyang, et al.
Pubblicazione: (2023)
Advanced Long-term Earth System Forecasting
di: Wu, Hao, et al.
Pubblicazione: (2025)
di: Wu, Hao, et al.
Pubblicazione: (2025)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
di: Wang, Chaoyi, et al.
Pubblicazione: (2025)
di: Wang, Chaoyi, et al.
Pubblicazione: (2025)
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
di: Li, Zhi, et al.
Pubblicazione: (2025)
di: Li, Zhi, et al.
Pubblicazione: (2025)
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
di: Wang, Fusang, et al.
Pubblicazione: (2026)
di: Wang, Fusang, et al.
Pubblicazione: (2026)
NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp Detection
di: Fan, Qingyu, et al.
Pubblicazione: (2025)
di: Fan, Qingyu, et al.
Pubblicazione: (2025)
DrivAer Transformer: A high-precision and fast prediction method for vehicle aerodynamic drag coefficient based on the DrivAerNet++ dataset
di: He, Jiaqi, et al.
Pubblicazione: (2025)
di: He, Jiaqi, et al.
Pubblicazione: (2025)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
di: Bian, Zhipeng, et al.
Pubblicazione: (2025)
di: Bian, Zhipeng, et al.
Pubblicazione: (2025)
Technology prediction of a 3D model using Neural Network
di: Miebs, Grzegorz, et al.
Pubblicazione: (2025)
di: Miebs, Grzegorz, et al.
Pubblicazione: (2025)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
di: Chen, Zhangquan, et al.
Pubblicazione: (2026)
di: Chen, Zhangquan, et al.
Pubblicazione: (2026)
A Segmented Robot Grasping Perception Neural Network for Edge AI
di: Bröcheler, Casper, et al.
Pubblicazione: (2025)
di: Bröcheler, Casper, et al.
Pubblicazione: (2025)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
di: Li, Yayuan, et al.
Pubblicazione: (2025)
di: Li, Yayuan, et al.
Pubblicazione: (2025)
Fast 3D point clouds retrieval for Large-scale 3D Place Recognition
di: Zede, Chahine-Nicolas, et al.
Pubblicazione: (2025)
di: Zede, Chahine-Nicolas, et al.
Pubblicazione: (2025)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
Flexible-weighted Chamfer Distance: Enhanced Objective Function for Point Cloud Completion
di: Li, Jie, et al.
Pubblicazione: (2025)
di: Li, Jie, et al.
Pubblicazione: (2025)
MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
di: Fan, Qiannan, et al.
Pubblicazione: (2025)
di: Fan, Qiannan, et al.
Pubblicazione: (2025)
Exploring Surround-View Fisheye Camera 3D Object Detection
di: Li, Changcai, et al.
Pubblicazione: (2025)
di: Li, Changcai, et al.
Pubblicazione: (2025)
Topology-Aware Latent Diffusion for 3D Shape Generation
di: Hu, Jiangbei, et al.
Pubblicazione: (2024)
di: Hu, Jiangbei, et al.
Pubblicazione: (2024)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
di: He, Jianxiang, et al.
Pubblicazione: (2025)
di: He, Jianxiang, et al.
Pubblicazione: (2025)
Hierarchical Feature-level Reverse Propagation for Post-Training Neural Networks
di: Ding, Ni, et al.
Pubblicazione: (2025)
di: Ding, Ni, et al.
Pubblicazione: (2025)
MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
di: Beilharz, Benjamin, et al.
Pubblicazione: (2025)
di: Beilharz, Benjamin, et al.
Pubblicazione: (2025)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
vS-Graphs: Tightly Coupling Visual SLAM and 3D Scene Graphs Exploiting Hierarchical Scene Understanding
di: Tourani, Ali, et al.
Pubblicazione: (2025)
di: Tourani, Ali, et al.
Pubblicazione: (2025)
A Simple Baseline for Streaming Video Understanding
di: Shen, Yujiao, et al.
Pubblicazione: (2026)
di: Shen, Yujiao, et al.
Pubblicazione: (2026)
Parking Space Detection in the City of Granada
di: Luis, Crespo-Orti, et al.
Pubblicazione: (2025)
di: Luis, Crespo-Orti, et al.
Pubblicazione: (2025)
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
di: Hu, Chenggong, et al.
Pubblicazione: (2026)
di: Hu, Chenggong, et al.
Pubblicazione: (2026)
Defending against Backdoor Attacks via Module Switching
di: Li, Weijun, et al.
Pubblicazione: (2025)
di: Li, Weijun, et al.
Pubblicazione: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
di: Su, Yuetong, et al.
Pubblicazione: (2025)
di: Su, Yuetong, et al.
Pubblicazione: (2025)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
di: Qian, Wenxu, et al.
Pubblicazione: (2025)
di: Qian, Wenxu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
di: Huang, Keli, et al.
Pubblicazione: (2022) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
di: Chen, Zhangquan, et al.
Pubblicazione: (2025) -
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
di: Han, Yudong, et al.
Pubblicazione: (2024) -
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
di: Mehta, Vinit, et al.
Pubblicazione: (2025)