VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Nonghai, Zhang, Zeyu, Wang, Jiazi, Zhao, Yang, Tang, Hao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
di: Ge, Jinchao, et al.
Pubblicazione: (2025)
di: Ge, Jinchao, et al.
Pubblicazione: (2025)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
Text-to-Image Synthesis: A Decade Survey
di: Zhang, Nonghai, et al.
Pubblicazione: (2024)
di: Zhang, Nonghai, et al.
Pubblicazione: (2024)
3D CoCa: Contrastive Learners are 3D Captioners
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
DragMesh: Interactive 3D Generation Made Easy
di: Zhang, Tianshan, et al.
Pubblicazione: (2025)
di: Zhang, Tianshan, et al.
Pubblicazione: (2025)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
di: Li, Peize, et al.
Pubblicazione: (2026)
di: Li, Peize, et al.
Pubblicazione: (2026)
MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing
di: Sun, Xiaokun, et al.
Pubblicazione: (2026)
di: Sun, Xiaokun, et al.
Pubblicazione: (2026)
UniMesh: Unifying 3D Mesh Understanding and Generation
di: Huang, Peng, et al.
Pubblicazione: (2026)
di: Huang, Peng, et al.
Pubblicazione: (2026)
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
di: Chen, Yixiong, et al.
Pubblicazione: (2026)
di: Chen, Yixiong, et al.
Pubblicazione: (2026)
3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence
di: Tang, Hao, et al.
Pubblicazione: (2026)
di: Tang, Hao, et al.
Pubblicazione: (2026)
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
Learn 3D VQA Better with Active Selection and Reannotation
di: Zhou, Shengli, et al.
Pubblicazione: (2025)
di: Zhou, Shengli, et al.
Pubblicazione: (2025)
Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning
di: Madhu, Prathmesh, et al.
Pubblicazione: (2020)
di: Madhu, Prathmesh, et al.
Pubblicazione: (2020)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
di: Zhang, Weichen, et al.
Pubblicazione: (2025)
di: Zhang, Weichen, et al.
Pubblicazione: (2025)
Code2Worlds: Empowering Coding LLMs for 4D World Generation
di: Zhang, Yi, et al.
Pubblicazione: (2026)
di: Zhang, Yi, et al.
Pubblicazione: (2026)
HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
di: Zhou, Shengli, et al.
Pubblicazione: (2025)
di: Zhou, Shengli, et al.
Pubblicazione: (2025)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
di: Liu, Yibo, et al.
Pubblicazione: (2024)
di: Liu, Yibo, et al.
Pubblicazione: (2024)
SplatTalk: 3D VQA with Gaussian Splatting
di: Thai, Anh, et al.
Pubblicazione: (2025)
di: Thai, Anh, et al.
Pubblicazione: (2025)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
di: Liu, Xinhang, et al.
Pubblicazione: (2025)
di: Liu, Xinhang, et al.
Pubblicazione: (2025)
Light4D: Training-Free Extreme Viewpoint 4D Video Relighting
di: Wu, Zhenghuang, et al.
Pubblicazione: (2026)
di: Wu, Zhenghuang, et al.
Pubblicazione: (2026)
Tetrahedron Splatting for 3D Generation
di: Gu, Chun, et al.
Pubblicazione: (2024)
di: Gu, Chun, et al.
Pubblicazione: (2024)
Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
di: Mansour, Adnan Ben, et al.
Pubblicazione: (2025)
di: Mansour, Adnan Ben, et al.
Pubblicazione: (2025)
UAVFF3D: A Geometry-Aware Benchmark for Feed-Forward UAV 3D Reconstruction
di: Yang, Xiang, et al.
Pubblicazione: (2026)
di: Yang, Xiang, et al.
Pubblicazione: (2026)
Efficient4D: Fast Dynamic 3D Object Generation from a Single-view Video
di: Pan, Zijie, et al.
Pubblicazione: (2024)
di: Pan, Zijie, et al.
Pubblicazione: (2024)
3D Primitives are a Spatial Language for VLMs
di: Liu, Junze, et al.
Pubblicazione: (2026)
di: Liu, Junze, et al.
Pubblicazione: (2026)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
di: He, Jixuan, et al.
Pubblicazione: (2026)
di: He, Jixuan, et al.
Pubblicazione: (2026)
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
di: Monon, Mashrafi, et al.
Pubblicazione: (2026)
di: Monon, Mashrafi, et al.
Pubblicazione: (2026)
Open-Pose 3D Zero-Shot Learning: Benchmark and Challenges
di: Zhao, Weiguang, et al.
Pubblicazione: (2023)
di: Zhao, Weiguang, et al.
Pubblicazione: (2023)
Multimodal LLM With Hierarchical Mixture-of-Experts for VQA on 3D Brain MRI
di: Vepa, Arvind Murari, et al.
Pubblicazione: (2025)
di: Vepa, Arvind Murari, et al.
Pubblicazione: (2025)
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
di: Cong, Wenyan, et al.
Pubblicazione: (2025)
di: Cong, Wenyan, et al.
Pubblicazione: (2025)
Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training
di: Lu, Hexiao, et al.
Pubblicazione: (2026)
di: Lu, Hexiao, et al.
Pubblicazione: (2026)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
di: Ma, Dongsheng, et al.
Pubblicazione: (2026)
di: Ma, Dongsheng, et al.
Pubblicazione: (2026)
MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets
di: Li, Zeyu, et al.
Pubblicazione: (2024)
di: Li, Zeyu, et al.
Pubblicazione: (2024)
Pixal3D: Pixel-Aligned 3D Generation from Images
di: Li, Dong-Yang, et al.
Pubblicazione: (2026)
di: Li, Dong-Yang, et al.
Pubblicazione: (2026)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment
di: Wang, Bin, et al.
Pubblicazione: (2025)
di: Wang, Bin, et al.
Pubblicazione: (2025)
LAA3D: A Benchmark of Detecting and Tracking Low-Altitude Aircraft in 3D Space
di: Wu, Hai, et al.
Pubblicazione: (2025)
di: Wu, Hai, et al.
Pubblicazione: (2025)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
di: Ge, Jinchao, et al.
Pubblicazione: (2025) -
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025) -
Text-to-Image Synthesis: A Decade Survey
di: Zhang, Nonghai, et al.
Pubblicazione: (2024) -
3D CoCa: Contrastive Learners are 3D Captioners
di: Huang, Ting, et al.
Pubblicazione: (2025) -
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)