Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiuming, Ni, Chaojun, Liu, Mengmeng, Peng, Chensheng, Wang, Fangjinhua, Shen, Sitian, Pollefeys, Marc, Tomizuka, Masayoshi, Tewari, Ayush, Kristensson, Per Ola |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias
by: Xu, Zhiyuan, et al.
Published: (2026)
by: Xu, Zhiyuan, et al.
Published: (2026)
Unbounded: Object-Boundary Interaction in Mixed Reality
by: Lyu, Zhuoyue, et al.
Published: (2025)
by: Lyu, Zhuoyue, et al.
Published: (2025)
Objestures: Everyday Objects Meet Mid-Air Gestures for Expressive Interaction
by: Lyu, Zhuoyue, et al.
Published: (2025)
by: Lyu, Zhuoyue, et al.
Published: (2025)
Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
by: Chen, Junlong, et al.
Published: (2024)
by: Chen, Junlong, et al.
Published: (2024)
Bend It, Aim It, Tap It: Designing an On-Body Disambiguation Mechanism for Curve Selection in Mixed Reality
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
The Argument is the Explanation: Structured Argumentation for Trust in Agents
by: Cakar, Ege, et al.
Published: (2025)
by: Cakar, Ege, et al.
Published: (2025)
Optimizing Curve-Based Selection with On-Body Surfaces in Virtual Environments
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Demystifying Reward Design in Reinforcement Learning for Upper Extremity Interaction: Practical Guidelines for Biomechanical Simulations in HCI
by: Selder, Hannah, et al.
Published: (2025)
by: Selder, Hannah, et al.
Published: (2025)
Generative AI for Accessible and Inclusive Extended Reality
by: Grubert, Jens, et al.
Published: (2024)
by: Grubert, Jens, et al.
Published: (2024)
How Do We Evaluate Experiences in Immersive Environments?
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Evaluating the Usability of Microgestures for Text Editing Tasks in Virtual Reality
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation
by: Yang, Boyin, et al.
Published: (2025)
by: Yang, Boyin, et al.
Published: (2025)
Large Language Model-assisted Speech and Pointing Benefits Multiple 3D Object Selection in Virtual Reality
by: Chen, Junlong, et al.
Published: (2024)
by: Chen, Junlong, et al.
Published: (2024)
Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception
by: Shen, Junxiao, et al.
Published: (2023)
by: Shen, Junxiao, et al.
Published: (2023)
MegaFlow: Zero-Shot Large Displacement Optical Flow
by: Zhang, Dingxi, et al.
Published: (2026)
by: Zhang, Dingxi, et al.
Published: (2026)
Nonparametric Inverse Dynamic Models for Multimodal Interactive Robots
by: Haninger, Kevin, et al.
Published: (2019)
by: Haninger, Kevin, et al.
Published: (2019)
RT-GS: Gaussian Splatting with Reflection and Transmittance Primitives
by: Zeng, Kunnong, et al.
Published: (2026)
by: Zeng, Kunnong, et al.
Published: (2026)
Joint Pedestrian Trajectory Prediction through Posterior Sampling
by: Lin, Haotian, et al.
Published: (2024)
by: Lin, Haotian, et al.
Published: (2024)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
Cost-Aware Bayesian Optimization for Prototyping Interactive Devices
by: Langerak, Thomas, et al.
Published: (2026)
by: Langerak, Thomas, et al.
Published: (2026)
Lightweight and Accurate Multi-View Stereo with Confidence-Aware Diffusion Model
by: Wang, Fangjinhua, et al.
Published: (2025)
by: Wang, Fangjinhua, et al.
Published: (2025)
Handows: A Palm-Based Interactive Multi-Window Management System in Virtual Reality
by: Wang, Jindu, et al.
Published: (2025)
by: Wang, Jindu, et al.
Published: (2025)
MyoInteract: A Framework for Fast Prototyping of Biomechanical HCI Tasks using Reinforcement Learning
by: Bhattarai, Ankit, et al.
Published: (2026)
by: Bhattarai, Ankit, et al.
Published: (2026)
Prompt-Driven Agentic Video Editing System: Autonomous Comprehension of Long-Form, Story-Driven Media
by: Ding, Zihan, et al.
Published: (2025)
by: Ding, Zihan, et al.
Published: (2025)
A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision
by: Peng, Chensheng, et al.
Published: (2024)
by: Peng, Chensheng, et al.
Published: (2024)
Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic Segmentation
by: Zheng, Yu, et al.
Published: (2023)
by: Zheng, Yu, et al.
Published: (2023)
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
by: Deng, Tianchen, et al.
Published: (2026)
by: Deng, Tianchen, et al.
Published: (2026)
A Multi-Camera Optical Tag Neuronavigation and AR Augmentation Framework for Non-Invasive Brain Stimulation
by: Hu, Xuyi, et al.
Published: (2026)
by: Hu, Xuyi, et al.
Published: (2026)
Optical Tag-Based Neuronavigation and Augmentation System for Non-Invasive Brain Stimulation
by: Hu, Xuyi, et al.
Published: (2026)
by: Hu, Xuyi, et al.
Published: (2026)
GLACE: Global Local Accelerated Coordinate Encoding
by: Wang, Fangjinhua, et al.
Published: (2024)
by: Wang, Fangjinhua, et al.
Published: (2024)
ImLoc: Revisiting Visual Localization with Image-based Representation
by: Jiang, Xudong, et al.
Published: (2026)
by: Jiang, Xudong, et al.
Published: (2026)
R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization
by: Jiang, Xudong, et al.
Published: (2025)
by: Jiang, Xudong, et al.
Published: (2025)
What Makes a Model Breathe? Understanding Reinforcement Learning Reward Function Design in Biomechanical User Simulation
by: Selder, Hannah, et al.
Published: (2025)
by: Selder, Hannah, et al.
Published: (2025)
Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces
by: Zhang, Chenyangguang, et al.
Published: (2025)
by: Zhang, Chenyangguang, et al.
Published: (2025)
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
by: Zhang, Chenyangguang, et al.
Published: (2026)
by: Zhang, Chenyangguang, et al.
Published: (2026)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
by: Zhou, Wenqi, et al.
Published: (2025)
by: Zhou, Wenqi, et al.
Published: (2025)
RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space
by: Xie, Yichen, et al.
Published: (2026)
by: Xie, Yichen, et al.
Published: (2026)
LocoScooter: Designing a Stationary Scooter-Based Locomotion System for Navigation in Virtual Reality
by: He, Wei, et al.
Published: (2026)
by: He, Wei, et al.
Published: (2026)
Adaptive Linear Path Model-Based Diffusion
by: Shimizu, Yutaka, et al.
Published: (2026)
by: Shimizu, Yutaka, et al.
Published: (2026)
Algebraic Control: Complete Stable Inversion with Necessary and Sufficient Conditions
by: Kürkçü, Burak, et al.
Published: (2024)
by: Kürkçü, Burak, et al.
Published: (2024)
Similar Items
-
Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias
by: Xu, Zhiyuan, et al.
Published: (2026) -
Unbounded: Object-Boundary Interaction in Mixed Reality
by: Lyu, Zhuoyue, et al.
Published: (2025) -
Objestures: Everyday Objects Meet Mid-Air Gestures for Expressive Interaction
by: Lyu, Zhuoyue, et al.
Published: (2025) -
Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
by: Chen, Junlong, et al.
Published: (2024) -
Bend It, Aim It, Tap It: Designing an On-Body Disambiguation Mechanism for Curve Selection in Mixed Reality
by: Li, Xiang, et al.
Published: (2025)