LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Yinglin, Zou, Zhengxia, Gu, Tongwei, Jia, Wei, Zhao, Zhan, Xu, Luyi, Liu, Xinzhu, Lin, Yenan, Jiang, Hao, Chen, Kang, Qiu, Shuang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WorldGPT: Empowering LLM as Multimodal World Model
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
MetaEarth3D: Unlocking World-scale 3D Generation with Spatially Scalable Generative Modeling
von: Cao, Jinqi, et al.
Veröffentlicht: (2026)
von: Cao, Jinqi, et al.
Veröffentlicht: (2026)
TriDF: Triplane-Accelerated Density Fields for Few-Shot Remote Sensing Novel View Synthesis
von: Kang, Jiaming, et al.
Veröffentlicht: (2025)
von: Kang, Jiaming, et al.
Veröffentlicht: (2025)
Code2Worlds: Empowering Coding LLMs for 4D World Generation
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
von: Li, Manyu, et al.
Veröffentlicht: (2026)
von: Li, Manyu, et al.
Veröffentlicht: (2026)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
BlockGaussian: Efficient Large-Scale Scene Novel View Synthesis via Adaptive Block-Based Gaussian Splatting
von: Wu, Yongchang, et al.
Veröffentlicht: (2025)
von: Wu, Yongchang, et al.
Veröffentlicht: (2025)
ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection
von: Sun, Zhihao, et al.
Veröffentlicht: (2024)
von: Sun, Zhihao, et al.
Veröffentlicht: (2024)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
von: Guo, Zile, et al.
Veröffentlicht: (2026)
von: Guo, Zile, et al.
Veröffentlicht: (2026)
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
von: Zhou, Runjie, et al.
Veröffentlicht: (2026)
von: Zhou, Runjie, et al.
Veröffentlicht: (2026)
The Moderating Effect of Informal Institutions: Clans and Straw Burning in China
von: Liang Tang, et al.
Veröffentlicht: (2025)
von: Liang Tang, et al.
Veröffentlicht: (2025)
The Complex and Challenging World of the Host–Pathogen Interaction
von: Marcel I. Ramirez
Veröffentlicht: (2024)
von: Marcel I. Ramirez
Veröffentlicht: (2024)
Unbiased Dynamic Multimodal Fusion
von: Wei, Shicai, et al.
Veröffentlicht: (2026)
von: Wei, Shicai, et al.
Veröffentlicht: (2026)
BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2025)
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2025)
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
von: Qiu, Wenmo, et al.
Veröffentlicht: (2024)
von: Qiu, Wenmo, et al.
Veröffentlicht: (2024)
Synergistic double‐doped elastic composites for durable, ultra‐flexible sign language translation sensors
von: Tongshun Wu, et al.
Veröffentlicht: (2025)
von: Tongshun Wu, et al.
Veröffentlicht: (2025)
SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations
von: Wu, Jason, et al.
Veröffentlicht: (2026)
von: Wu, Jason, et al.
Veröffentlicht: (2026)
Efficient Semantic Splatting for Remote Sensing Multi-view Segmentation
von: Qi, Zipeng, et al.
Veröffentlicht: (2024)
von: Qi, Zipeng, et al.
Veröffentlicht: (2024)
Change-Agent: Towards Interactive Comprehensive Remote Sensing Change Interpretation and Analysis
von: Liu, Chenyang, et al.
Veröffentlicht: (2024)
von: Liu, Chenyang, et al.
Veröffentlicht: (2024)
Empowering Multi-Robot Cooperation via Sequential World Models
von: Zhao, Zijie, et al.
Veröffentlicht: (2025)
von: Zhao, Zijie, et al.
Veröffentlicht: (2025)
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
von: Li, Teng, et al.
Veröffentlicht: (2025)
von: Li, Teng, et al.
Veröffentlicht: (2025)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2025)
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
von: GigaWorld Team, et al.
Veröffentlicht: (2025)
von: GigaWorld Team, et al.
Veröffentlicht: (2025)
A Performance Investigation of Multimodal Multiobjective Optimization Algorithms in Solving Two Types of Real-World Problems
von: Chen, Zhiqiu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiqiu, et al.
Veröffentlicht: (2024)
Matrix-Game: Interactive World Foundation Model
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Spatial-Temporal Human-Object Interaction Detection
von: Sun, Xu, et al.
Veröffentlicht: (2025)
von: Sun, Xu, et al.
Veröffentlicht: (2025)
Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
von: Zheng, Guangze, et al.
Veröffentlicht: (2025)
von: Zheng, Guangze, et al.
Veröffentlicht: (2025)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
Olaf-World: Orienting Latent Actions for Video World Modeling
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
Empowering Teachers to Build a Better World
Veröffentlicht: (2020)
Veröffentlicht: (2020)
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
von: Guo, Junliang, et al.
Veröffentlicht: (2025)
von: Guo, Junliang, et al.
Veröffentlicht: (2025)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
von: Xu, Xiaojie, et al.
Veröffentlicht: (2026)
von: Xu, Xiaojie, et al.
Veröffentlicht: (2026)
Large AI Model Empowered Multimodal Semantic Communications
von: Jiang, Feibo, et al.
Veröffentlicht: (2023)
von: Jiang, Feibo, et al.
Veröffentlicht: (2023)
Can Large Language Models Understand Real-World Complex Instructions?
von: He, Qianyu, et al.
Veröffentlicht: (2023)
von: He, Qianyu, et al.
Veröffentlicht: (2023)
The DAWN of World-Action Interactive Models
von: Lu, Hongbo, et al.
Veröffentlicht: (2026)
von: Lu, Hongbo, et al.
Veröffentlicht: (2026)
WSSM: Geographic-enhanced hierarchical state-space model for global station weather forecast
von: Yang, Songru, et al.
Veröffentlicht: (2025)
von: Yang, Songru, et al.
Veröffentlicht: (2025)
CloudMamba: An Uncertainty-Guided Dual-Scale Mamba Network for Cloud Detection in Remote Sensing Imagery
von: Yang, Jiajun, et al.
Veröffentlicht: (2026)
von: Yang, Jiajun, et al.
Veröffentlicht: (2026)
On Large Multimodal Models as Open-World Image Classifiers
von: Conti, Alessandro, et al.
Veröffentlicht: (2025)
von: Conti, Alessandro, et al.
Veröffentlicht: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
von: Zhang, Haichao, et al.
Veröffentlicht: (2026)
von: Zhang, Haichao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
WorldGPT: Empowering LLM as Multimodal World Model
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024) -
MetaEarth3D: Unlocking World-scale 3D Generation with Spatially Scalable Generative Modeling
von: Cao, Jinqi, et al.
Veröffentlicht: (2026) -
TriDF: Triplane-Accelerated Density Fields for Few-Shot Remote Sensing Novel View Synthesis
von: Kang, Jiaming, et al.
Veröffentlicht: (2025) -
Code2Worlds: Empowering Coding LLMs for 4D World Generation
von: Zhang, Yi, et al.
Veröffentlicht: (2026) -
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
von: Li, Manyu, et al.
Veröffentlicht: (2026)