Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Shangwen, Peng, Qianyu, Pu, Zhao, Shu, Zhilei, Ke, Xiangrui, Xing, Zhaohu, Tong, Zizhao, Wang, Zeqing, Cui, Xinyu, Wang, Huangji, Zhao, Jian, Jin, Yeying, Cheng, Fan, Feng, Ruili |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
by: Tong, Zizhao, et al.
Published: (2026)
by: Tong, Zizhao, et al.
Published: (2026)
ReactiveGWM: Steering NPC in Reactive Game World Models
by: Wang, Zeqing, et al.
Published: (2026)
by: Wang, Zeqing, et al.
Published: (2026)
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
by: Zhu, Shangwen, et al.
Published: (2025)
by: Zhu, Shangwen, et al.
Published: (2025)
TIE: Time Interval Encoding for Video Generation over Events
by: Shu, Zhilei, et al.
Published: (2026)
by: Shu, Zhilei, et al.
Published: (2026)
Addressing the ID-Matching Challenge in Long Video Captioning
by: Yang, Zhantao, et al.
Published: (2025)
by: Yang, Zhantao, et al.
Published: (2025)
RAIN: Real-time Animation of Infinite Video Stream
by: Shu, Zhilei, et al.
Published: (2024)
by: Shu, Zhilei, et al.
Published: (2024)
Leveraging Large Language Models for Entity Matching
by: Huang, Qianyu, et al.
Published: (2024)
by: Huang, Qianyu, et al.
Published: (2024)
Healing Magic and Evil Demons. Canonical Udug-hul Incantations
by: Geller, Markham J.
Published: (2019)
by: Geller, Markham J.
Published: (2019)
Entity-Guided Multi-Task Learning for Infrared and Visible Image Fusion
by: Shao, Wenyu, et al.
Published: (2026)
by: Shao, Wenyu, et al.
Published: (2026)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
Dyn-O: Building Structured World Models with Object-Centric Representations
by: Wang, Zizhao, et al.
Published: (2025)
by: Wang, Zizhao, et al.
Published: (2025)
Factored Latent Action World Models
by: Wang, Zizhao, et al.
Published: (2026)
by: Wang, Zizhao, et al.
Published: (2026)
Nahm sum identities for Cartan matrices of type $D_k$
by: Wang, Liuquan, et al.
Published: (2025)
by: Wang, Liuquan, et al.
Published: (2025)
The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control
by: Feng, Ruili, et al.
Published: (2024)
by: Feng, Ruili, et al.
Published: (2024)
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
by: Yang, Zhantao, et al.
Published: (2024)
by: Yang, Zhantao, et al.
Published: (2024)
Toward Real-World High-Precision Image Matting and Segmentation
by: Zhou, Haipeng, et al.
Published: (2026)
by: Zhou, Haipeng, et al.
Published: (2026)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
by: Yu, Chang, et al.
Published: (2024)
by: Yu, Chang, et al.
Published: (2024)
Phosphonoalamides Reveal the Biosynthetic Origin of Phosphonoalanine Natural Products and a Convergent Pathway for Their Diversification
by: Jerry J. Cui, et al.
Published: (2024)
by: Jerry J. Cui, et al.
Published: (2024)
UAU-Net: Uncertainty-aware Representation Learning and Evidential Classification for Facial Action Unit Detection
by: Li, Yuze, et al.
Published: (2026)
by: Li, Yuze, et al.
Published: (2026)
Giant Optical Nonlinear Response up to 60th‐Order Induced by the Ytterbium Energy Relay Mediated Photon Avalanches
by: Chenyi Wang, et al.
Published: (2024)
by: Chenyi Wang, et al.
Published: (2024)
Cusps and boundaries of connected fundamental domains for $Γ_0(N)$
by: Nie, Zhaohu
Published: (2025)
by: Nie, Zhaohu
Published: (2025)
Blowup masses of Toda systems corresponding to the Weyl groups
by: Nie, Zhaohu
Published: (2025)
by: Nie, Zhaohu
Published: (2025)
Instability in Diffusion ODEs: An Explanation for Inaccurate Image Reconstruction
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image Segmentation
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
Natural gamma radiation of ODP Hole 184-1148A
by: Tian, Jun, et al.
Published: (2008)
by: Tian, Jun, et al.
Published: (2008)
Adaptive soft sensor modeling of chemical processes based on an improved just‐in‐time learning and random mapping partial least squares
by: Ke Zhang, et al.
Published: (2024)
by: Ke Zhang, et al.
Published: (2024)
Teacher-Student Network for Real-World Face Super-Resolution with Progressive Embedding of Edge Information
by: Liu, Zhilei, et al.
Published: (2024)
by: Liu, Zhilei, et al.
Published: (2024)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
by: Chen, Sen, et al.
Published: (2022)
by: Chen, Sen, et al.
Published: (2022)
Targeted Regulation of Plant Organ Morphogenesis via Nanocarrier‐Based miRNA Delivery Systems
by: Bingyao Jiang, et al.
Published: (2025)
by: Bingyao Jiang, et al.
Published: (2025)
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
by: Wu, Shu, et al.
Published: (2026)
by: Wu, Shu, et al.
Published: (2026)
Emergent Continuous Time Crystal in Dissipative Quantum Spin System without Driving
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
Rabi Oscillation of High Partial Wave Interacting Atoms in Deep Optical Lattice
by: Wang, Zeqing, et al.
Published: (2024)
by: Wang, Zeqing, et al.
Published: (2024)
Single Particle Spectroscopies of $p$-wave and $d$-wave Interacting Bose Gases in Normal Phase
by: Wang, Zeqing, et al.
Published: (2023)
by: Wang, Zeqing, et al.
Published: (2023)
Effect of ultrasonic modification on the binding ability of pectin to anthocyanin
by: Yutong Liu, et al.
Published: (2024)
by: Yutong Liu, et al.
Published: (2024)
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
by: Ye, Angen, et al.
Published: (2026)
by: Ye, Angen, et al.
Published: (2026)
ActRef: Enhancing the Understanding of Python Code Refactoring with Action-Based Analysis
by: Wang, Siqi, et al.
Published: (2025)
by: Wang, Siqi, et al.
Published: (2025)
LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution
by: Fei, Song, et al.
Published: (2026)
by: Fei, Song, et al.
Published: (2026)
Diff-VPS: Video Polyp Segmentation via a Multi-task Diffusion Network with Adversarial Temporal Reasoning
by: Lu, Yingling, et al.
Published: (2024)
by: Lu, Yingling, et al.
Published: (2024)
Co-Evolving Latent Action World Models
by: Wang, Yucen, et al.
Published: (2025)
by: Wang, Yucen, et al.
Published: (2025)
Similar Items
-
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
by: Tong, Zizhao, et al.
Published: (2026) -
ReactiveGWM: Steering NPC in Reactive Game World Models
by: Wang, Zeqing, et al.
Published: (2026) -
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
by: Zhu, Shangwen, et al.
Published: (2025) -
TIE: Time Interval Encoding for Video Generation over Events
by: Shu, Zhilei, et al.
Published: (2026) -
Addressing the ID-Matching Challenge in Long Video Captioning
by: Yang, Zhantao, et al.
Published: (2025)