WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Runjie, Shao, Youbo, Lu, Haoyu, Xing, Bowei, Bai, Tongtong, Chen, Yujie, Zhao, Jie, Sui, Lin, Yao, Haotian, Zhao, Zijia, Yang, Hao, Wu, Haoning, Zhou, Zaida, Zhu, Jinguo, Huang, Zhiqi, Bao, Yiping, Liu, Yangyang, Charles, Y., Zhou, Xinyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Pixel-Level VLM Perception via Simple Points Prediction
di: Song, Tianhui, et al.
Pubblicazione: (2026)
di: Song, Tianhui, et al.
Pubblicazione: (2026)
CBPNet: A Continual Backpropagation Prompt Network for Alleviating Plasticity Loss on Edge Devices
di: Shao, Runjie, et al.
Pubblicazione: (2025)
di: Shao, Runjie, et al.
Pubblicazione: (2025)
WorldGPT: Empowering LLM as Multimodal World Model
di: Ge, Zhiqi, et al.
Pubblicazione: (2024)
di: Ge, Zhiqi, et al.
Pubblicazione: (2024)
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
di: Guo, Yangyang, et al.
Pubblicazione: (2023)
di: Guo, Yangyang, et al.
Pubblicazione: (2023)
Neuro-Symbolic Synergy for Interactive World Modeling
di: Zhao, Hongyu, et al.
Pubblicazione: (2026)
di: Zhao, Hongyu, et al.
Pubblicazione: (2026)
Semantic-Guided Dynamic Sparsification for Pre-Trained Model-based Class-Incremental Learning
di: Liu, Ruiqi, et al.
Pubblicazione: (2026)
di: Liu, Ruiqi, et al.
Pubblicazione: (2026)
DynaDrag: Dynamic Drag-Style Image Editing by Motion Prediction
di: Sui, Jiacheng, et al.
Pubblicazione: (2026)
di: Sui, Jiacheng, et al.
Pubblicazione: (2026)
Boosting General Trimap-free Matting in the Real-World Image
di: Zhao, Leo Shan Wenzhang Zhou Grace
Pubblicazione: (2024)
di: Zhao, Leo Shan Wenzhang Zhou Grace
Pubblicazione: (2024)
Toward an Integrated Cross-Urban Accident Prevention System: A Multi-Task Spatial-Temporal Learning Framework for Urban Safety Management
di: Fang, Jiayu, et al.
Pubblicazione: (2026)
di: Fang, Jiayu, et al.
Pubblicazione: (2026)
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
di: Zhou, Xunchu, et al.
Pubblicazione: (2024)
di: Zhou, Xunchu, et al.
Pubblicazione: (2024)
WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems
di: Wang, Yuchen, et al.
Pubblicazione: (2026)
di: Wang, Yuchen, et al.
Pubblicazione: (2026)
Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion
di: Wei, Shuoyan, et al.
Pubblicazione: (2026)
di: Wei, Shuoyan, et al.
Pubblicazione: (2026)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
di: Ma, Jiatong, et al.
Pubblicazione: (2026)
di: Ma, Jiatong, et al.
Pubblicazione: (2026)
Hybrid Gauge Approach for Accurate Real-Time TDDFT Simulations with Numerical Atomic Orbitals
di: Zhao, Haotian, et al.
Pubblicazione: (2025)
di: Zhao, Haotian, et al.
Pubblicazione: (2025)
An Oxygen‐Self‐Produced Nanoplatform Based on MnO2 for Relieving Hypoxia
di: Guohua Pan, et al.
Pubblicazione: (2024)
di: Guohua Pan, et al.
Pubblicazione: (2024)
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
di: Huang, Jing, et al.
Pubblicazione: (2025)
di: Huang, Jing, et al.
Pubblicazione: (2025)
An Empirical Study of World Model Quantization
di: Fu, Zhongqian, et al.
Pubblicazione: (2026)
di: Fu, Zhongqian, et al.
Pubblicazione: (2026)
WorldCompass: Reinforcement Learning for Long-Horizon World Models
di: Wang, Zehan, et al.
Pubblicazione: (2026)
di: Wang, Zehan, et al.
Pubblicazione: (2026)
DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels
di: Guo, Erjian, et al.
Pubblicazione: (2025)
di: Guo, Erjian, et al.
Pubblicazione: (2025)
User Prompting Strategies and ChatGPT Contextual Adaptation Shape Conversational Information-Seeking Experiences
di: Xue, Haoning, et al.
Pubblicazione: (2025)
di: Xue, Haoning, et al.
Pubblicazione: (2025)
Active Neural Mapping at Scale
di: Kuang, Zijia, et al.
Pubblicazione: (2024)
di: Kuang, Zijia, et al.
Pubblicazione: (2024)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
di: Zhou, Xin, et al.
Pubblicazione: (2026)
di: Zhou, Xin, et al.
Pubblicazione: (2026)
WorldSimBench: Towards Video Generation Models as World Simulators
di: Qin, Yiran, et al.
Pubblicazione: (2024)
di: Qin, Yiran, et al.
Pubblicazione: (2024)
Homogeneous linear recurrence relations of the determinants of distance matrices of trees
di: Liu, Zhiqi, et al.
Pubblicazione: (2025)
di: Liu, Zhiqi, et al.
Pubblicazione: (2025)
Precise Deviations for the Ewens-Pitman Model
di: Peng, Zhiqi, et al.
Pubblicazione: (2025)
di: Peng, Zhiqi, et al.
Pubblicazione: (2025)
Premium sharing for supply chain coordination with business interruption insurance under supply disruption
di: Yingjian Lei, et al.
Pubblicazione: (2025)
di: Yingjian Lei, et al.
Pubblicazione: (2025)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
di: Liu, Yuan, et al.
Pubblicazione: (2024)
di: Liu, Yuan, et al.
Pubblicazione: (2024)
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
di: Wang, Wenxuan, et al.
Pubblicazione: (2024)
Aperiodic intermittent containment consensus control for uncertain multi‐agent systems based on disturbance observer and input saturation
di: Beining Bao, et al.
Pubblicazione: (2025)
di: Beining Bao, et al.
Pubblicazione: (2025)
HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
di: Song, Quanjian, et al.
Pubblicazione: (2025)
di: Song, Quanjian, et al.
Pubblicazione: (2025)
Strong Lattice Softening Induced by Atomic Mismatch in Meta‐Phase Thermoelectrics
di: Kunpeng Zhao, et al.
Pubblicazione: (2025)
di: Kunpeng Zhao, et al.
Pubblicazione: (2025)
The Real‐World Analysis of Adverse Events With Two Types of Single‐Inhaler Triple Therapy: A VigiAccess Database Study
di: Baiquan Zhang, et al.
Pubblicazione: (2025)
di: Baiquan Zhang, et al.
Pubblicazione: (2025)
Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark
di: Zhou, Hua, et al.
Pubblicazione: (2025)
di: Zhou, Hua, et al.
Pubblicazione: (2025)
A homotopical consequence of branched covers
di: Hu, Runjie
Pubblicazione: (2023)
di: Hu, Runjie
Pubblicazione: (2023)
$L$-theory Characteristic Classes
di: Hu, Runjie
Pubblicazione: (2023)
di: Hu, Runjie
Pubblicazione: (2023)
Galois Symmetry of $Gal(\overline{\mathbb{Q}}/\mathbb{Q})$ on Topological Manifold Structures of Varieties
di: Hu, Runjie
Pubblicazione: (2023)
di: Hu, Runjie
Pubblicazione: (2023)
TesserAct: Learning 4D Embodied World Models
di: Zhen, Haoyu, et al.
Pubblicazione: (2025)
di: Zhen, Haoyu, et al.
Pubblicazione: (2025)
REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
di: Shi, Changyue, et al.
Pubblicazione: (2025)
di: Shi, Changyue, et al.
Pubblicazione: (2025)
Disentanglement-Based Equivariant Learning for Compositional VQA
di: Du, Zhou, et al.
Pubblicazione: (2026)
di: Du, Zhou, et al.
Pubblicazione: (2026)
Does Parental Migration Alleviate Multidimensional Poverty among Left‐behind Children in Rural Areas?
di: Yexin Zhou, et al.
Pubblicazione: (2025)
di: Yexin Zhou, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards Pixel-Level VLM Perception via Simple Points Prediction
di: Song, Tianhui, et al.
Pubblicazione: (2026) -
CBPNet: A Continual Backpropagation Prompt Network for Alleviating Plasticity Loss on Edge Devices
di: Shao, Runjie, et al.
Pubblicazione: (2025) -
WorldGPT: Empowering LLM as Multimodal World Model
di: Ge, Zhiqi, et al.
Pubblicazione: (2024) -
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
di: Guo, Yangyang, et al.
Pubblicazione: (2023) -
Neuro-Symbolic Synergy for Interactive World Modeling
di: Zhao, Hongyu, et al.
Pubblicazione: (2026)