WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Runjie, Shao, Youbo, Lu, Haoyu, Xing, Bowei, Bai, Tongtong, Chen, Yujie, Zhao, Jie, Sui, Lin, Yao, Haotian, Zhao, Zijia, Yang, Hao, Wu, Haoning, Zhou, Zaida, Zhu, Jinguo, Huang, Zhiqi, Bao, Yiping, Liu, Yangyang, Charles, Y., Zhou, Xinyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Pixel-Level VLM Perception via Simple Points Prediction
por: Song, Tianhui, et al.
Publicado: (2026)
por: Song, Tianhui, et al.
Publicado: (2026)
CBPNet: A Continual Backpropagation Prompt Network for Alleviating Plasticity Loss on Edge Devices
por: Shao, Runjie, et al.
Publicado: (2025)
por: Shao, Runjie, et al.
Publicado: (2025)
WorldGPT: Empowering LLM as Multimodal World Model
por: Ge, Zhiqi, et al.
Publicado: (2024)
por: Ge, Zhiqi, et al.
Publicado: (2024)
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
por: Guo, Yangyang, et al.
Publicado: (2023)
por: Guo, Yangyang, et al.
Publicado: (2023)
Neuro-Symbolic Synergy for Interactive World Modeling
por: Zhao, Hongyu, et al.
Publicado: (2026)
por: Zhao, Hongyu, et al.
Publicado: (2026)
Semantic-Guided Dynamic Sparsification for Pre-Trained Model-based Class-Incremental Learning
por: Liu, Ruiqi, et al.
Publicado: (2026)
por: Liu, Ruiqi, et al.
Publicado: (2026)
DynaDrag: Dynamic Drag-Style Image Editing by Motion Prediction
por: Sui, Jiacheng, et al.
Publicado: (2026)
por: Sui, Jiacheng, et al.
Publicado: (2026)
Boosting General Trimap-free Matting in the Real-World Image
por: Zhao, Leo Shan Wenzhang Zhou Grace
Publicado: (2024)
por: Zhao, Leo Shan Wenzhang Zhou Grace
Publicado: (2024)
Toward an Integrated Cross-Urban Accident Prevention System: A Multi-Task Spatial-Temporal Learning Framework for Urban Safety Management
por: Fang, Jiayu, et al.
Publicado: (2026)
por: Fang, Jiayu, et al.
Publicado: (2026)
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
por: Zhou, Xunchu, et al.
Publicado: (2024)
por: Zhou, Xunchu, et al.
Publicado: (2024)
WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems
por: Wang, Yuchen, et al.
Publicado: (2026)
por: Wang, Yuchen, et al.
Publicado: (2026)
Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion
por: Wei, Shuoyan, et al.
Publicado: (2026)
por: Wei, Shuoyan, et al.
Publicado: (2026)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
por: Ma, Jiatong, et al.
Publicado: (2026)
por: Ma, Jiatong, et al.
Publicado: (2026)
Hybrid Gauge Approach for Accurate Real-Time TDDFT Simulations with Numerical Atomic Orbitals
por: Zhao, Haotian, et al.
Publicado: (2025)
por: Zhao, Haotian, et al.
Publicado: (2025)
An Oxygen‐Self‐Produced Nanoplatform Based on MnO2 for Relieving Hypoxia
por: Guohua Pan, et al.
Publicado: (2024)
por: Guohua Pan, et al.
Publicado: (2024)
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
por: Huang, Jing, et al.
Publicado: (2025)
por: Huang, Jing, et al.
Publicado: (2025)
An Empirical Study of World Model Quantization
por: Fu, Zhongqian, et al.
Publicado: (2026)
por: Fu, Zhongqian, et al.
Publicado: (2026)
WorldCompass: Reinforcement Learning for Long-Horizon World Models
por: Wang, Zehan, et al.
Publicado: (2026)
por: Wang, Zehan, et al.
Publicado: (2026)
DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels
por: Guo, Erjian, et al.
Publicado: (2025)
por: Guo, Erjian, et al.
Publicado: (2025)
User Prompting Strategies and ChatGPT Contextual Adaptation Shape Conversational Information-Seeking Experiences
por: Xue, Haoning, et al.
Publicado: (2025)
por: Xue, Haoning, et al.
Publicado: (2025)
Active Neural Mapping at Scale
por: Kuang, Zijia, et al.
Publicado: (2024)
por: Kuang, Zijia, et al.
Publicado: (2024)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
por: Zhou, Xin, et al.
Publicado: (2026)
por: Zhou, Xin, et al.
Publicado: (2026)
WorldSimBench: Towards Video Generation Models as World Simulators
por: Qin, Yiran, et al.
Publicado: (2024)
por: Qin, Yiran, et al.
Publicado: (2024)
Homogeneous linear recurrence relations of the determinants of distance matrices of trees
por: Liu, Zhiqi, et al.
Publicado: (2025)
por: Liu, Zhiqi, et al.
Publicado: (2025)
Precise Deviations for the Ewens-Pitman Model
por: Peng, Zhiqi, et al.
Publicado: (2025)
por: Peng, Zhiqi, et al.
Publicado: (2025)
Premium sharing for supply chain coordination with business interruption insurance under supply disruption
por: Yingjian Lei, et al.
Publicado: (2025)
por: Yingjian Lei, et al.
Publicado: (2025)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
por: Liu, Yuan, et al.
Publicado: (2024)
por: Liu, Yuan, et al.
Publicado: (2024)
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions
por: Wang, Wenxuan, et al.
Publicado: (2024)
por: Wang, Wenxuan, et al.
Publicado: (2024)
Aperiodic intermittent containment consensus control for uncertain multi‐agent systems based on disturbance observer and input saturation
por: Beining Bao, et al.
Publicado: (2025)
por: Beining Bao, et al.
Publicado: (2025)
HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
por: Song, Quanjian, et al.
Publicado: (2025)
por: Song, Quanjian, et al.
Publicado: (2025)
Strong Lattice Softening Induced by Atomic Mismatch in Meta‐Phase Thermoelectrics
por: Kunpeng Zhao, et al.
Publicado: (2025)
por: Kunpeng Zhao, et al.
Publicado: (2025)
The Real‐World Analysis of Adverse Events With Two Types of Single‐Inhaler Triple Therapy: A VigiAccess Database Study
por: Baiquan Zhang, et al.
Publicado: (2025)
por: Baiquan Zhang, et al.
Publicado: (2025)
Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark
por: Zhou, Hua, et al.
Publicado: (2025)
por: Zhou, Hua, et al.
Publicado: (2025)
A homotopical consequence of branched covers
por: Hu, Runjie
Publicado: (2023)
por: Hu, Runjie
Publicado: (2023)
$L$-theory Characteristic Classes
por: Hu, Runjie
Publicado: (2023)
por: Hu, Runjie
Publicado: (2023)
Galois Symmetry of $Gal(\overline{\mathbb{Q}}/\mathbb{Q})$ on Topological Manifold Structures of Varieties
por: Hu, Runjie
Publicado: (2023)
por: Hu, Runjie
Publicado: (2023)
TesserAct: Learning 4D Embodied World Models
por: Zhen, Haoyu, et al.
Publicado: (2025)
por: Zhen, Haoyu, et al.
Publicado: (2025)
REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
por: Shi, Changyue, et al.
Publicado: (2025)
por: Shi, Changyue, et al.
Publicado: (2025)
Disentanglement-Based Equivariant Learning for Compositional VQA
por: Du, Zhou, et al.
Publicado: (2026)
por: Du, Zhou, et al.
Publicado: (2026)
Does Parental Migration Alleviate Multidimensional Poverty among Left‐behind Children in Rural Areas?
por: Yexin Zhou, et al.
Publicado: (2025)
por: Yexin Zhou, et al.
Publicado: (2025)
Ejemplares similares
-
Towards Pixel-Level VLM Perception via Simple Points Prediction
por: Song, Tianhui, et al.
Publicado: (2026) -
CBPNet: A Continual Backpropagation Prompt Network for Alleviating Plasticity Loss on Edge Devices
por: Shao, Runjie, et al.
Publicado: (2025) -
WorldGPT: Empowering LLM as Multimodal World Model
por: Ge, Zhiqi, et al.
Publicado: (2024) -
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
por: Guo, Yangyang, et al.
Publicado: (2023) -
Neuro-Symbolic Synergy for Interactive World Modeling
por: Zhao, Hongyu, et al.
Publicado: (2026)