IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Parker, Li, Chenxin, Li, Zhengxin, Wu, Yipeng, Li, Wuyang, Yang, Zhiqin, Zhang, Zhenyuan, Lin, Yunlong, Han, Sirui, Feng, Brandon Y. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LGS: A Light-weight 4D Gaussian Splatting for Efficient Surgical Scene Reconstruction
di: Liu, Hengyu, et al.
Pubblicazione: (2024)
di: Liu, Hengyu, et al.
Pubblicazione: (2024)
TensoIR: Tensorial Inverse Rendering
di: Jin, Haian, et al.
Pubblicazione: (2023)
di: Jin, Haian, et al.
Pubblicazione: (2023)
UrbanIR: Large-Scale Urban Scene Inverse Rendering from a Single Video
di: Lin, Chih-Hao, et al.
Pubblicazione: (2023)
di: Lin, Chih-Hao, et al.
Pubblicazione: (2023)
GUS-IR: Gaussian Splatting with Unified Shading for Inverse Rendering
di: Liang, Zhihao, et al.
Pubblicazione: (2024)
di: Liang, Zhihao, et al.
Pubblicazione: (2024)
SAR-GS: Gaussian Splatting based SAR Images Rendering and Target Reconstruction
di: Li, Aobo, et al.
Pubblicazione: (2025)
di: Li, Aobo, et al.
Pubblicazione: (2025)
GS-IR: 3D Gaussian Splatting for Inverse Rendering
di: Liang, Zhihao, et al.
Pubblicazione: (2023)
di: Liang, Zhihao, et al.
Pubblicazione: (2023)
JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
di: Lin, Yunlong, et al.
Pubblicazione: (2025)
di: Lin, Yunlong, et al.
Pubblicazione: (2025)
IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range Images
di: Lin, Chih-Hao, et al.
Pubblicazione: (2024)
di: Lin, Chih-Hao, et al.
Pubblicazione: (2024)
DiffuSyn Bench: Evaluating Vision-Language Models on Real-World Complexities with Diffusion-Generated Synthetic Benchmarks
di: Zhou, Haokun, et al.
Pubblicazione: (2024)
di: Zhou, Haokun, et al.
Pubblicazione: (2024)
Endora: Video Generation Models as Endoscopy Simulators
di: Li, Chenxin, et al.
Pubblicazione: (2024)
di: Li, Chenxin, et al.
Pubblicazione: (2024)
MemFly: On-the-Fly Memory Optimization via Information Bottleneck
di: Zhang, Zhenyuan, et al.
Pubblicazione: (2026)
di: Zhang, Zhenyuan, et al.
Pubblicazione: (2026)
EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian Splatting
di: Li, Chenxin, et al.
Pubblicazione: (2024)
di: Li, Chenxin, et al.
Pubblicazione: (2024)
Harnessing Lightweight Transformer with Contextual Synergic Enhancement for Efficient 3D Medical Image Segmentation
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
di: Ling, Lu, et al.
Pubblicazione: (2025)
di: Ling, Lu, et al.
Pubblicazione: (2025)
SVG-IR: Spatially-Varying Gaussian Splatting for Inverse Rendering
di: Sun, Hanxiao, et al.
Pubblicazione: (2025)
di: Sun, Hanxiao, et al.
Pubblicazione: (2025)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
di: Li, Lei, et al.
Pubblicazione: (2024)
di: Li, Lei, et al.
Pubblicazione: (2024)
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity
di: Wei, Xuefeng, et al.
Pubblicazione: (2026)
di: Wei, Xuefeng, et al.
Pubblicazione: (2026)
GIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization
di: Shi, Yahao, et al.
Pubblicazione: (2023)
di: Shi, Yahao, et al.
Pubblicazione: (2023)
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
di: Gui, Rui, et al.
Pubblicazione: (2025)
di: Gui, Rui, et al.
Pubblicazione: (2025)
LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and Control
di: Qu, Delin, et al.
Pubblicazione: (2024)
di: Qu, Delin, et al.
Pubblicazione: (2024)
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
di: Li, Yue, et al.
Pubblicazione: (2025)
di: Li, Yue, et al.
Pubblicazione: (2025)
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
di: Xie, Qinghongbing, et al.
Pubblicazione: (2025)
di: Xie, Qinghongbing, et al.
Pubblicazione: (2025)
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
di: Pan, Panwang, et al.
Pubblicazione: (2025)
di: Pan, Panwang, et al.
Pubblicazione: (2025)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
di: Lin, Jingru, et al.
Pubblicazione: (2025)
di: Lin, Jingru, et al.
Pubblicazione: (2025)
AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
di: Wang, Yonghui, et al.
Pubblicazione: (2024)
di: Wang, Yonghui, et al.
Pubblicazione: (2024)
RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
di: Xiao, Xi, et al.
Pubblicazione: (2025)
di: Xiao, Xi, et al.
Pubblicazione: (2025)
FedGPS: Statistical Rectification Against Data Heterogeneity in Federated Learning
di: Yang, Zhiqin, et al.
Pubblicazione: (2025)
di: Yang, Zhiqin, et al.
Pubblicazione: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
di: Lin, Jingli, et al.
Pubblicazione: (2025)
di: Lin, Jingli, et al.
Pubblicazione: (2025)
WildCap: Facial Albedo Capture in the Wild via Hybrid Inverse Rendering
di: Han, Yuxuan, et al.
Pubblicazione: (2025)
di: Han, Yuxuan, et al.
Pubblicazione: (2025)
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
di: Duan, Lin, et al.
Pubblicazione: (2025)
di: Duan, Lin, et al.
Pubblicazione: (2025)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
di: Jia, Baoxiong, et al.
Pubblicazione: (2024)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
di: Tian, Kexin, et al.
Pubblicazione: (2025)
di: Tian, Kexin, et al.
Pubblicazione: (2025)
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
di: Huang, Yuzhi, et al.
Pubblicazione: (2025)
di: Huang, Yuzhi, et al.
Pubblicazione: (2025)
EventVL: Understand Event Streams via Multimodal Large Language Model
di: Li, Pengteng, et al.
Pubblicazione: (2025)
di: Li, Pengteng, et al.
Pubblicazione: (2025)
Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenes
di: Choi, JunYong, et al.
Pubblicazione: (2025)
di: Choi, JunYong, et al.
Pubblicazione: (2025)
UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation
di: Wu, Shuang, et al.
Pubblicazione: (2024)
di: Wu, Shuang, et al.
Pubblicazione: (2024)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
di: Chow, Wei, et al.
Pubblicazione: (2025)
di: Chow, Wei, et al.
Pubblicazione: (2025)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
di: Yu, Haorui, et al.
Pubblicazione: (2026)
di: Yu, Haorui, et al.
Pubblicazione: (2026)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
di: Gao, Haoxiang, et al.
Pubblicazione: (2025)
di: Gao, Haoxiang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LGS: A Light-weight 4D Gaussian Splatting for Efficient Surgical Scene Reconstruction
di: Liu, Hengyu, et al.
Pubblicazione: (2024) -
TensoIR: Tensorial Inverse Rendering
di: Jin, Haian, et al.
Pubblicazione: (2023) -
UrbanIR: Large-Scale Urban Scene Inverse Rendering from a Single Video
di: Lin, Chih-Hao, et al.
Pubblicazione: (2023) -
GUS-IR: Gaussian Splatting with Unified Shading for Inverse Rendering
di: Liang, Zhihao, et al.
Pubblicazione: (2024) -
SAR-GS: Gaussian Splatting based SAR Images Rendering and Target Reconstruction
di: Li, Aobo, et al.
Pubblicazione: (2025)