Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Shangwen, Peng, Qianyu, Pu, Zhao, Shu, Zhilei, Ke, Xiangrui, Xing, Zhaohu, Tong, Zizhao, Wang, Zeqing, Cui, Xinyu, Wang, Huangji, Zhao, Jian, Jin, Yeying, Cheng, Fan, Feng, Ruili |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
von: Tong, Zizhao, et al.
Veröffentlicht: (2026)
von: Tong, Zizhao, et al.
Veröffentlicht: (2026)
ReactiveGWM: Steering NPC in Reactive Game World Models
von: Wang, Zeqing, et al.
Veröffentlicht: (2026)
von: Wang, Zeqing, et al.
Veröffentlicht: (2026)
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
von: Zhu, Shangwen, et al.
Veröffentlicht: (2025)
von: Zhu, Shangwen, et al.
Veröffentlicht: (2025)
TIE: Time Interval Encoding for Video Generation over Events
von: Shu, Zhilei, et al.
Veröffentlicht: (2026)
von: Shu, Zhilei, et al.
Veröffentlicht: (2026)
Addressing the ID-Matching Challenge in Long Video Captioning
von: Yang, Zhantao, et al.
Veröffentlicht: (2025)
von: Yang, Zhantao, et al.
Veröffentlicht: (2025)
RAIN: Real-time Animation of Infinite Video Stream
von: Shu, Zhilei, et al.
Veröffentlicht: (2024)
von: Shu, Zhilei, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Entity Matching
von: Huang, Qianyu, et al.
Veröffentlicht: (2024)
von: Huang, Qianyu, et al.
Veröffentlicht: (2024)
Healing Magic and Evil Demons. Canonical Udug-hul Incantations
von: Geller, Markham J.
Veröffentlicht: (2019)
von: Geller, Markham J.
Veröffentlicht: (2019)
Entity-Guided Multi-Task Learning for Infrared and Visible Image Fusion
von: Shao, Wenyu, et al.
Veröffentlicht: (2026)
von: Shao, Wenyu, et al.
Veröffentlicht: (2026)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
Dyn-O: Building Structured World Models with Object-Centric Representations
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
Factored Latent Action World Models
von: Wang, Zizhao, et al.
Veröffentlicht: (2026)
von: Wang, Zizhao, et al.
Veröffentlicht: (2026)
Nahm sum identities for Cartan matrices of type $D_k$
von: Wang, Liuquan, et al.
Veröffentlicht: (2025)
von: Wang, Liuquan, et al.
Veröffentlicht: (2025)
The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control
von: Feng, Ruili, et al.
Veröffentlicht: (2024)
von: Feng, Ruili, et al.
Veröffentlicht: (2024)
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
Toward Real-World High-Precision Image Matting and Segmentation
von: Zhou, Haipeng, et al.
Veröffentlicht: (2026)
von: Zhou, Haipeng, et al.
Veröffentlicht: (2026)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
von: Yu, Chang, et al.
Veröffentlicht: (2024)
von: Yu, Chang, et al.
Veröffentlicht: (2024)
Phosphonoalamides Reveal the Biosynthetic Origin of Phosphonoalanine Natural Products and a Convergent Pathway for Their Diversification
von: Jerry J. Cui, et al.
Veröffentlicht: (2024)
von: Jerry J. Cui, et al.
Veröffentlicht: (2024)
UAU-Net: Uncertainty-aware Representation Learning and Evidential Classification for Facial Action Unit Detection
von: Li, Yuze, et al.
Veröffentlicht: (2026)
von: Li, Yuze, et al.
Veröffentlicht: (2026)
Giant Optical Nonlinear Response up to 60th‐Order Induced by the Ytterbium Energy Relay Mediated Photon Avalanches
von: Chenyi Wang, et al.
Veröffentlicht: (2024)
von: Chenyi Wang, et al.
Veröffentlicht: (2024)
Cusps and boundaries of connected fundamental domains for $Γ_0(N)$
von: Nie, Zhaohu
Veröffentlicht: (2025)
von: Nie, Zhaohu
Veröffentlicht: (2025)
Blowup masses of Toda systems corresponding to the Weyl groups
von: Nie, Zhaohu
Veröffentlicht: (2025)
von: Nie, Zhaohu
Veröffentlicht: (2025)
Instability in Diffusion ODEs: An Explanation for Inaccurate Image Reconstruction
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image Segmentation
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
Natural gamma radiation of ODP Hole 184-1148A
von: Tian, Jun, et al.
Veröffentlicht: (2008)
von: Tian, Jun, et al.
Veröffentlicht: (2008)
Adaptive soft sensor modeling of chemical processes based on an improved just‐in‐time learning and random mapping partial least squares
von: Ke Zhang, et al.
Veröffentlicht: (2024)
von: Ke Zhang, et al.
Veröffentlicht: (2024)
Teacher-Student Network for Real-World Face Super-Resolution with Progressive Embedding of Edge Information
von: Liu, Zhilei, et al.
Veröffentlicht: (2024)
von: Liu, Zhilei, et al.
Veröffentlicht: (2024)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
von: Chen, Sen, et al.
Veröffentlicht: (2022)
von: Chen, Sen, et al.
Veröffentlicht: (2022)
Targeted Regulation of Plant Organ Morphogenesis via Nanocarrier‐Based miRNA Delivery Systems
von: Bingyao Jiang, et al.
Veröffentlicht: (2025)
von: Bingyao Jiang, et al.
Veröffentlicht: (2025)
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
von: Wu, Shu, et al.
Veröffentlicht: (2026)
von: Wu, Shu, et al.
Veröffentlicht: (2026)
Emergent Continuous Time Crystal in Dissipative Quantum Spin System without Driving
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
von: Wang, Zeqing, et al.
Veröffentlicht: (2025)
Rabi Oscillation of High Partial Wave Interacting Atoms in Deep Optical Lattice
von: Wang, Zeqing, et al.
Veröffentlicht: (2024)
von: Wang, Zeqing, et al.
Veröffentlicht: (2024)
Single Particle Spectroscopies of $p$-wave and $d$-wave Interacting Bose Gases in Normal Phase
von: Wang, Zeqing, et al.
Veröffentlicht: (2023)
von: Wang, Zeqing, et al.
Veröffentlicht: (2023)
Effect of ultrasonic modification on the binding ability of pectin to anthocyanin
von: Yutong Liu, et al.
Veröffentlicht: (2024)
von: Yutong Liu, et al.
Veröffentlicht: (2024)
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
von: Ye, Angen, et al.
Veröffentlicht: (2026)
von: Ye, Angen, et al.
Veröffentlicht: (2026)
ActRef: Enhancing the Understanding of Python Code Refactoring with Action-Based Analysis
von: Wang, Siqi, et al.
Veröffentlicht: (2025)
von: Wang, Siqi, et al.
Veröffentlicht: (2025)
LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution
von: Fei, Song, et al.
Veröffentlicht: (2026)
von: Fei, Song, et al.
Veröffentlicht: (2026)
Diff-VPS: Video Polyp Segmentation via a Multi-task Diffusion Network with Adversarial Temporal Reasoning
von: Lu, Yingling, et al.
Veröffentlicht: (2024)
von: Lu, Yingling, et al.
Veröffentlicht: (2024)
Co-Evolving Latent Action World Models
von: Wang, Yucen, et al.
Veröffentlicht: (2025)
von: Wang, Yucen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
von: Tong, Zizhao, et al.
Veröffentlicht: (2026) -
ReactiveGWM: Steering NPC in Reactive Game World Models
von: Wang, Zeqing, et al.
Veröffentlicht: (2026) -
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
von: Zhu, Shangwen, et al.
Veröffentlicht: (2025) -
TIE: Time Interval Encoding for Video Generation over Events
von: Shu, Zhilei, et al.
Veröffentlicht: (2026) -
Addressing the ID-Matching Challenge in Long Video Captioning
von: Yang, Zhantao, et al.
Veröffentlicht: (2025)