Pandora: Towards General World Model with Natural Language Actions and Video States
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xiang, Jiannan, Liu, Guangyi, Gu, Yi, Gao, Qiyue, Ning, Yuting, Zha, Yuheng, Feng, Zeyu, Tao, Tianhua, Hao, Shibo, Shi, Yemin, Liu, Zhengzhong, Xing, Eric P., Hu, Zhiting |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
World Reasoning Arena
par: PAN Team, et autres
Publié: (2026)
par: PAN Team, et autres
Publié: (2026)
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
par: Zha, Yuheng, et autres
Publié: (2025)
par: Zha, Yuheng, et autres
Publié: (2025)
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
par: Shi, Yemin, et autres
Publié: (2025)
par: Shi, Yemin, et autres
Publié: (2025)
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
par: PAN Team, et autres
Publié: (2025)
par: PAN Team, et autres
Publié: (2025)
Critiques of World Models
par: Xing, Eric, et autres
Publié: (2025)
par: Xing, Eric, et autres
Publié: (2025)
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
par: Hao, Shibo, et autres
Publié: (2023)
par: Hao, Shibo, et autres
Publié: (2023)
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
par: Yin, Yanbin, et autres
Publié: (2025)
par: Yin, Yanbin, et autres
Publié: (2025)
General Agentic Planning Through Simulative Reasoning with World Models
par: Deng, Mingkai, et autres
Publié: (2025)
par: Deng, Mingkai, et autres
Publié: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
par: Liu, Chengzhi, et autres
Publié: (2026)
par: Liu, Chengzhi, et autres
Publié: (2026)
SlimPajama-DC: Understanding Data Combinations for LLM Training
par: Shen, Zhiqiang, et autres
Publié: (2023)
par: Shen, Zhiqiang, et autres
Publié: (2023)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
par: Hao, Shibo, et autres
Publié: (2024)
par: Hao, Shibo, et autres
Publié: (2024)
3D CoCa: Contrastive Learners are 3D Captioners
par: Huang, Ting, et autres
Publié: (2025)
par: Huang, Ting, et autres
Publié: (2025)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
par: Zhu, Shangwen, et autres
Publié: (2026)
par: Zhu, Shangwen, et autres
Publié: (2026)
VSA: Faster Video Diffusion with Trainable Sparse Attention
par: Zhang, Peiyuan, et autres
Publié: (2025)
par: Zhang, Peiyuan, et autres
Publié: (2025)
MWM: Mobile World Models for Action-Conditioned Consistent Prediction
par: Yan, Han, et autres
Publié: (2026)
par: Yan, Han, et autres
Publié: (2026)
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
par: Gao, Qiyue, et autres
Publié: (2025)
par: Gao, Qiyue, et autres
Publié: (2025)
How Confident are Video Models? Empowering Video Models to Express their Uncertainty
par: Mei, Zhiting, et autres
Publié: (2025)
par: Mei, Zhiting, et autres
Publié: (2025)
No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors
par: Liu, Tao, et autres
Publié: (2026)
par: Liu, Tao, et autres
Publié: (2026)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
par: Li, Huiqiong, et autres
Publié: (2026)
par: Li, Huiqiong, et autres
Publié: (2026)
The DAWN of World-Action Interactive Models
par: Lu, Hongbo, et autres
Publié: (2026)
par: Lu, Hongbo, et autres
Publié: (2026)
LangCoop: Collaborative Driving with Language
par: Gao, Xiangbo, et autres
Publié: (2025)
par: Gao, Xiangbo, et autres
Publié: (2025)
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
par: Cheng, Zhoujun, et autres
Publié: (2025)
par: Cheng, Zhoujun, et autres
Publié: (2025)
ANVIL: Accelerator-Native Video Interpolation via Codec Motion Vector Priors
par: Liu, Shibo
Publié: (2026)
par: Liu, Shibo
Publié: (2026)
Crystal: Illuminating LLM Abilities on Language and Code
par: Tao, Tianhua, et autres
Publié: (2024)
par: Tao, Tianhua, et autres
Publié: (2024)
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models
par: Singla, Somanshu, et autres
Publié: (2024)
par: Singla, Somanshu, et autres
Publié: (2024)
Olaf-World: Orienting Latent Actions for Video World Modeling
par: Jiang, Yuxin, et autres
Publié: (2026)
par: Jiang, Yuxin, et autres
Publié: (2026)
Token Level Routing Inference System for Edge Devices
par: She, Jianshu, et autres
Publié: (2025)
par: She, Jianshu, et autres
Publié: (2025)
Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMs
par: Tan, Bowen, et autres
Publié: (2025)
par: Tan, Bowen, et autres
Publié: (2025)
VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
par: Ji, Yuheng, et autres
Publié: (2025)
par: Ji, Yuheng, et autres
Publié: (2025)
HiLight: Technical Report on the Motern AI Video Language Model
par: Wang, Zhiting, et autres
Publié: (2024)
par: Wang, Zhiting, et autres
Publié: (2024)
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
par: Yin, Tenny, et autres
Publié: (2025)
par: Yin, Tenny, et autres
Publié: (2025)
PISCO: Precise Video Instance Insertion with Sparse Control
par: Gao, Xiangbo, et autres
Publié: (2026)
par: Gao, Xiangbo, et autres
Publié: (2026)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
par: Wu, Yuheng, et autres
Publié: (2026)
par: Wu, Yuheng, et autres
Publié: (2026)
CocoaBench: Evaluating Unified Digital Agents in the Wild
par: CocoaBench Team, et autres
Publié: (2026)
par: CocoaBench Team, et autres
Publié: (2026)
Markovian Pandora's box
par: Yang, Yuanyuan, et autres
Publié: (2025)
par: Yang, Yuanyuan, et autres
Publié: (2025)
RTS-Mono: A Real-Time Self-Supervised Monocular Depth Estimation Method for Real-World Deployment
par: Cheng, Zeyu, et autres
Publié: (2025)
par: Cheng, Zeyu, et autres
Publié: (2025)
MultiHateLoc: Towards Temporal Localisation of Multimodal Hate Content in Online Videos
par: Sun, Qiyue, et autres
Publié: (2025)
par: Sun, Qiyue, et autres
Publié: (2025)
Global and Local Entailment Learning for Natural World Imagery
par: Sastry, Srikumar, et autres
Publié: (2025)
par: Sastry, Srikumar, et autres
Publié: (2025)
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
par: Mei, Zhiting, et autres
Publié: (2025)
par: Mei, Zhiting, et autres
Publié: (2025)
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language Models
par: Fan, Zhiting, et autres
Publié: (2025)
par: Fan, Zhiting, et autres
Publié: (2025)
Documents similaires
-
World Reasoning Arena
par: PAN Team, et autres
Publié: (2026) -
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
par: Zha, Yuheng, et autres
Publié: (2025) -
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
par: Shi, Yemin, et autres
Publié: (2025) -
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
par: PAN Team, et autres
Publié: (2025) -
Critiques of World Models
par: Xing, Eric, et autres
Publié: (2025)