SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Haosong, Li, Hao, Chen, Jiaqi, Pan, Yuhao, Yao, Runmao, Dai, Yalun, Huo, Fushuo, Hong, Fangzhou, Chen, Zhaoxi, Wang, Haozhao, Zhang, Dingwen, Liu, Ziwei, Xu, Wenchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
PrismWF: A Multi-Granularity Patch-Based Transformer for Robust Website Fingerprinting Attack
von: Pan, Yuhao, et al.
Veröffentlicht: (2026)
von: Pan, Yuhao, et al.
Veröffentlicht: (2026)
SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
von: Workman, Kenny, et al.
Veröffentlicht: (2025)
von: Workman, Kenny, et al.
Veröffentlicht: (2025)
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
von: Cao, Ziang, et al.
Veröffentlicht: (2026)
von: Cao, Ziang, et al.
Veröffentlicht: (2026)
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
von: Huo, Fushuo, et al.
Veröffentlicht: (2024)
von: Huo, Fushuo, et al.
Veröffentlicht: (2024)
Overcome Modal Bias in Multi-modal Federated Learning via Balanced Modality Selection
von: Fan, Yunfeng, et al.
Veröffentlicht: (2023)
von: Fan, Yunfeng, et al.
Veröffentlicht: (2023)
SpatialBench-UC: Uncertainty-Aware Evaluation of Spatial Prompt Following in Text-to-Image Generation
von: Rostane, Amine
Veröffentlicht: (2026)
von: Rostane, Amine
Veröffentlicht: (2026)
Cross-Receiver Generalization for RF Fingerprint Identification via Feature Disentanglement and Adversarial Training
von: Pan, Yuhao, et al.
Veröffentlicht: (2025)
von: Pan, Yuhao, et al.
Veröffentlicht: (2025)
Towards Robust Multimodal Learning in the Open World
von: Huo, Fushuo
Veröffentlicht: (2025)
von: Huo, Fushuo
Veröffentlicht: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
MMBench: Is Your Multi-modal Model an All-around Player?
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
von: Liu, Yuan, et al.
Veröffentlicht: (2023)
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
von: Xie, Haozhe, et al.
Veröffentlicht: (2023)
von: Xie, Haozhe, et al.
Veröffentlicht: (2023)
Generative Gaussian Splatting for Unbounded 3D City Generation
von: Xie, Haozhe, et al.
Veröffentlicht: (2024)
von: Xie, Haozhe, et al.
Veröffentlicht: (2024)
Compositional Generative Model of Unbounded 4D Cities
von: Xie, Haozhe, et al.
Veröffentlicht: (2025)
von: Xie, Haozhe, et al.
Veröffentlicht: (2025)
FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls
von: Hu, Tao, et al.
Veröffentlicht: (2024)
von: Hu, Tao, et al.
Veröffentlicht: (2024)
REQA: Coarse-to-fine Assessment of Image Quality to Alleviate the Range Effect
von: Li, Bingheng, et al.
Veröffentlicht: (2022)
von: Li, Bingheng, et al.
Veröffentlicht: (2022)
Reconstructing 4D Spatial Intelligence: A Survey
von: Cao, Yukang, et al.
Veröffentlicht: (2025)
von: Cao, Yukang, et al.
Veröffentlicht: (2025)
3D Scene Generation: A Survey
von: Wen, Beichen, et al.
Veröffentlicht: (2025)
von: Wen, Beichen, et al.
Veröffentlicht: (2025)
Is Your Driving World Model an All-Around Player?
von: Kong, Lingdong, et al.
Veröffentlicht: (2026)
von: Kong, Lingdong, et al.
Veröffentlicht: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
Radiant: Large-scale 3D Gaussian Rendering based on Hierarchical Framework
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
von: Cao, Yukang, et al.
Veröffentlicht: (2026)
von: Cao, Yukang, et al.
Veröffentlicht: (2026)
PhysX-3D: Physical-Grounded 3D Asset Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
von: Liu, Zixian, et al.
Veröffentlicht: (2026)
von: Liu, Zixian, et al.
Veröffentlicht: (2026)
AirRoom: Objects Matter in Room Reidentification
von: Yao, Runmao, et al.
Veröffentlicht: (2025)
von: Yao, Runmao, et al.
Veröffentlicht: (2025)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
von: Fan, Yunfeng, et al.
Veröffentlicht: (2023)
von: Fan, Yunfeng, et al.
Veröffentlicht: (2023)
Scaling Spatial Intelligence with Multimodal Foundation Models
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark
von: Wang, Pan, et al.
Veröffentlicht: (2025)
von: Wang, Pan, et al.
Veröffentlicht: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
von: Xie, Haozhe, et al.
Veröffentlicht: (2026)
von: Xie, Haozhe, et al.
Veröffentlicht: (2026)
PhilEO Bench: Evaluating Geo-Spatial Foundation Models
von: Fibaek, Casper, et al.
Veröffentlicht: (2024)
von: Fibaek, Casper, et al.
Veröffentlicht: (2024)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
PushupBench: Your VLM is not good at counting pushups
von: Li, Shengzhi, et al.
Veröffentlicht: (2026)
von: Li, Shengzhi, et al.
Veröffentlicht: (2026)
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
von: Chen, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoxi, et al.
Veröffentlicht: (2025)
Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
Research on the Spatial Data Intelligent Foundation Model
von: Wang, Shaohua, et al.
Veröffentlicht: (2024)
von: Wang, Shaohua, et al.
Veröffentlicht: (2024)
Efficient Deep Demosaicing with Spatially Downsampled Isotropic Networks
von: Fan, Cory, et al.
Veröffentlicht: (2026)
von: Fan, Cory, et al.
Veröffentlicht: (2026)
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
von: Zheng, Peirong, et al.
Veröffentlicht: (2026)
von: Zheng, Peirong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
von: Xu, Peiran, et al.
Veröffentlicht: (2025) -
PrismWF: A Multi-Granularity Patch-Based Transformer for Robust Website Fingerprinting Attack
von: Pan, Yuhao, et al.
Veröffentlicht: (2026) -
SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
von: Workman, Kenny, et al.
Veröffentlicht: (2025) -
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
von: Cao, Ziang, et al.
Veröffentlicht: (2026) -
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
von: Huo, Fushuo, et al.
Veröffentlicht: (2024)