Dense360: Dense Understanding from Omnidirectional Panoramas

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yikang, Zhang, Tao, Zhang, Dizhe, Ji, Shunping, Li, Xiangtai, Qi, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908410737328128
author Zhou, Yikang
Zhang, Tao
Zhang, Dizhe
Ji, Shunping
Li, Xiangtai
Qi, Lu
author_facet Zhou, Yikang
Zhang, Tao
Zhang, Dizhe
Ji, Shunping
Li, Xiangtai
Qi, Lu
contents Multimodal Large Language Models (MLLMs) require comprehensive visual inputs to achieve dense understanding of the physical world. While existing MLLMs demonstrate impressive world understanding capabilities through limited field-of-view (FOV) visual inputs (e.g., 70 degree), we take the first step toward dense understanding from omnidirectional panoramas. We first introduce an omnidirectional panoramas dataset featuring a comprehensive suite of reliability-scored annotations. Specifically, our dataset contains 160K panoramas with 5M dense entity-level captions, 1M unique referring expressions, and 100K entity-grounded panoramic scene descriptions. Compared to multi-view alternatives, panoramas can provide more complete, compact, and continuous scene representations through equirectangular projections (ERP). However, the use of ERP introduces two key challenges for MLLMs: i) spatial continuity along the circle of latitude, and ii) latitude-dependent variation in information density. We address these challenges through ERP-RoPE, a position encoding scheme specifically designed for panoramic ERP. In addition, we introduce Dense360-Bench, the first benchmark for evaluating MLLMs on omnidirectional captioning and grounding, establishing a comprehensive framework for advancing dense visual-language understanding in panoramic settings.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14471
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dense360: Dense Understanding from Omnidirectional Panoramas
Zhou, Yikang
Zhang, Tao
Zhang, Dizhe
Ji, Shunping
Li, Xiangtai
Qi, Lu
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) require comprehensive visual inputs to achieve dense understanding of the physical world. While existing MLLMs demonstrate impressive world understanding capabilities through limited field-of-view (FOV) visual inputs (e.g., 70 degree), we take the first step toward dense understanding from omnidirectional panoramas. We first introduce an omnidirectional panoramas dataset featuring a comprehensive suite of reliability-scored annotations. Specifically, our dataset contains 160K panoramas with 5M dense entity-level captions, 1M unique referring expressions, and 100K entity-grounded panoramic scene descriptions. Compared to multi-view alternatives, panoramas can provide more complete, compact, and continuous scene representations through equirectangular projections (ERP). However, the use of ERP introduces two key challenges for MLLMs: i) spatial continuity along the circle of latitude, and ii) latitude-dependent variation in information density. We address these challenges through ERP-RoPE, a position encoding scheme specifically designed for panoramic ERP. In addition, we introduce Dense360-Bench, the first benchmark for evaluating MLLMs on omnidirectional captioning and grounding, establishing a comprehensive framework for advancing dense visual-language understanding in panoramic settings.
title Dense360: Dense Understanding from Omnidirectional Panoramas
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.14471