GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Ge, Yongtao, Xu, Guangkai, Zhao, Zhiyue, Sun, Libo, Huang, Zheng, Sun, Yanlong, Chen, Hao, Shen, Chunhua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
di: Xu, Guangkai, et al.
Pubblicazione: (2026)
di: Xu, Guangkai, et al.
Pubblicazione: (2026)
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
di: Xu, Guangkai, et al.
Pubblicazione: (2024)
di: Xu, Guangkai, et al.
Pubblicazione: (2024)
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
di: Zhang, Songyan, et al.
Pubblicazione: (2025)
di: Zhang, Songyan, et al.
Pubblicazione: (2025)
Generative Video Matting
di: Ge, Yongtao, et al.
Pubblicazione: (2025)
di: Ge, Yongtao, et al.
Pubblicazione: (2025)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
Diffusion Models are Efficient Data Generators for Human Mesh Recovery
di: Ge, Yongtao, et al.
Pubblicazione: (2024)
di: Ge, Yongtao, et al.
Pubblicazione: (2024)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
di: Zhu, Muzhi, et al.
Pubblicazione: (2024)
di: Zhu, Muzhi, et al.
Pubblicazione: (2024)
DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation
di: He, Xiankang, et al.
Pubblicazione: (2024)
di: He, Xiankang, et al.
Pubblicazione: (2024)
Guided Diffusion-based Generation of Adversarial Objects for Real-World Monocular Depth Estimation Attacks
di: Chen, Yongtao, et al.
Pubblicazione: (2025)
di: Chen, Yongtao, et al.
Pubblicazione: (2025)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
di: Li, Liyang, et al.
Pubblicazione: (2026)
di: Li, Liyang, et al.
Pubblicazione: (2026)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
di: Li, Liyang, et al.
Pubblicazione: (2026)
di: Li, Liyang, et al.
Pubblicazione: (2026)
YOLO-NAS-Bench: A Surrogate Benchmark with Self-Evolving Predictors for YOLO Architecture Search
di: Li, Zhe, et al.
Pubblicazione: (2026)
di: Li, Zhe, et al.
Pubblicazione: (2026)
Physical 3D Adversarial Attacks against Monocular Depth Estimation in Autonomous Driving
di: Zheng, Junhao, et al.
Pubblicazione: (2024)
di: Zheng, Junhao, et al.
Pubblicazione: (2024)
Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation
di: Hu, Mu, et al.
Pubblicazione: (2024)
di: Hu, Mu, et al.
Pubblicazione: (2024)
Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation
di: Cai, Xinhao, et al.
Pubblicazione: (2026)
di: Cai, Xinhao, et al.
Pubblicazione: (2026)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
di: Li, Zizun, et al.
Pubblicazione: (2026)
di: Li, Zizun, et al.
Pubblicazione: (2026)
MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception
di: Hao, Xiaoshuai, et al.
Pubblicazione: (2025)
di: Hao, Xiaoshuai, et al.
Pubblicazione: (2025)
RGM: A Robust Generalizable Matching Model
di: Zhang, Songyan, et al.
Pubblicazione: (2023)
di: Zhang, Songyan, et al.
Pubblicazione: (2023)
GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
di: Zheng, Yushuo, et al.
Pubblicazione: (2025)
di: Zheng, Yushuo, et al.
Pubblicazione: (2025)
RAFT-MSF++: Temporal Geometry-Motion Feature Fusion for Self-Supervised Monocular Scene Flow
di: Sun, Xunpei, et al.
Pubblicazione: (2026)
di: Sun, Xunpei, et al.
Pubblicazione: (2026)
Towards Robust Monocular Depth Estimation in Non-Lambertian Surfaces
di: Zhang, Junrui, et al.
Pubblicazione: (2024)
di: Zhang, Junrui, et al.
Pubblicazione: (2024)
GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry
di: He, Xiankang, et al.
Pubblicazione: (2026)
di: He, Xiankang, et al.
Pubblicazione: (2026)
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
di: Sun, Yang-Tian, et al.
Pubblicazione: (2025)
di: Sun, Yang-Tian, et al.
Pubblicazione: (2025)
GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction
di: Lin, Weiquan, et al.
Pubblicazione: (2026)
di: Lin, Weiquan, et al.
Pubblicazione: (2026)
Adaptive Surface Normal Constraint for Geometric Estimation from Monocular Images
di: Long, Xiaoxiao, et al.
Pubblicazione: (2024)
di: Long, Xiaoxiao, et al.
Pubblicazione: (2024)
Benchmark on Monocular Metric Depth Estimation in Wildlife Setting
di: Niccoli, Niccolò, et al.
Pubblicazione: (2025)
di: Niccoli, Niccolò, et al.
Pubblicazione: (2025)
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
di: Luo, Luqing, et al.
Pubblicazione: (2024)
di: Luo, Luqing, et al.
Pubblicazione: (2024)
MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
di: Wang, Ruicheng, et al.
Pubblicazione: (2025)
di: Wang, Ruicheng, et al.
Pubblicazione: (2025)
Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
di: Park, Seong Hyeon, et al.
Pubblicazione: (2025)
di: Park, Seong Hyeon, et al.
Pubblicazione: (2025)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
di: Huang, Junming, et al.
Pubblicazione: (2026)
di: Huang, Junming, et al.
Pubblicazione: (2026)
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
di: Yin, Zijin, et al.
Pubblicazione: (2026)
di: Yin, Zijin, et al.
Pubblicazione: (2026)
MultiGO++: Monocular 3D Clothed Human Reconstruction via Geometry-Texture Collaboration
di: Yao, Nanjie, et al.
Pubblicazione: (2026)
di: Yao, Nanjie, et al.
Pubblicazione: (2026)
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
di: Gao, Tianyi, et al.
Pubblicazione: (2025)
di: Gao, Tianyi, et al.
Pubblicazione: (2025)
FlowDepth: Decoupling Optical Flow for Self-Supervised Monocular Depth Estimation
di: Sun, Yiyang, et al.
Pubblicazione: (2024)
di: Sun, Yiyang, et al.
Pubblicazione: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos
di: Rho, Daniel, et al.
Pubblicazione: (2026)
di: Rho, Daniel, et al.
Pubblicazione: (2026)
Geometry-Constrained Monocular Scale Estimation Using Semantic Segmentation for Dynamic Scenes
di: Zhang, Hui, et al.
Pubblicazione: (2025)
di: Zhang, Hui, et al.
Pubblicazione: (2025)
Real-time Monocular Depth Estimation on Embedded Systems
di: Feng, Cheng, et al.
Pubblicazione: (2023)
di: Feng, Cheng, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
di: Xu, Guangkai, et al.
Pubblicazione: (2026) -
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
di: Xu, Guangkai, et al.
Pubblicazione: (2024) -
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
di: Feng, Yuan, et al.
Pubblicazione: (2025) -
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
di: Zhao, Canyu, et al.
Pubblicazione: (2025) -
POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
di: Zhang, Songyan, et al.
Pubblicazione: (2025)