VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Zhongwei, Wei, Yunchao, Yu, Xiao, Luo, Guixun, Zhao, Yao, Kang, Bingyi, Feng, Jiashi, Jin, Xiaojie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
by: Ren, Zhongwei, et al.
Published: (2025)
by: Ren, Zhongwei, et al.
Published: (2025)
PixelLM: Pixel Reasoning with Large Multimodal Model
by: Ren, Zhongwei, et al.
Published: (2023)
by: Ren, Zhongwei, et al.
Published: (2023)
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
by: Xing, Ke, et al.
Published: (2025)
by: Xing, Ke, et al.
Published: (2025)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
by: Zhang, Haoji, et al.
Published: (2025)
by: Zhang, Haoji, et al.
Published: (2025)
How Far is Video Generation from World Model: A Physical Law Perspective
by: Kang, Bingyi, et al.
Published: (2024)
by: Kang, Bingyi, et al.
Published: (2024)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)
by: Chen, Sili, et al.
Published: (2025)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
by: Zhang, Haoji, et al.
Published: (2024)
by: Zhang, Haoji, et al.
Published: (2024)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
by: Yin, Yuyang, et al.
Published: (2025)
by: Yin, Yuyang, et al.
Published: (2025)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
Hierarchical Memory for Long Video QA
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
by: Ma, Fan, et al.
Published: (2023)
by: Ma, Fan, et al.
Published: (2023)
A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
by: Lin, Dongheng, et al.
Published: (2025)
by: Lin, Dongheng, et al.
Published: (2025)
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
Classification Done Right for Vision-Language Pre-Training
by: Huang, Zilong, et al.
Published: (2024)
by: Huang, Zilong, et al.
Published: (2024)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion
by: Wei, Shuoyan, et al.
Published: (2026)
by: Wei, Shuoyan, et al.
Published: (2026)
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
by: Xu, Xiaojie, et al.
Published: (2026)
by: Xu, Xiaojie, et al.
Published: (2026)
GenWorld: Towards Detecting AI-generated Real-world Simulation Videos
by: Chen, Weiliang, et al.
Published: (2025)
by: Chen, Weiliang, et al.
Published: (2025)
Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
by: Liang, Lili, et al.
Published: (2024)
by: Liang, Lili, et al.
Published: (2024)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
by: Xu, Lin, et al.
Published: (2024)
by: Xu, Lin, et al.
Published: (2024)
IPSeg: Image Posterior Mitigates Semantic Drift in Class-Incremental Segmentation
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
RealViformer: Investigating Attention for Real-World Video Super-Resolution
by: Zhang, Yuehan, et al.
Published: (2024)
by: Zhang, Yuehan, et al.
Published: (2024)
Transferable and Principled Efficiency for Open-Vocabulary Segmentation
by: Xu, Jingxuan, et al.
Published: (2024)
by: Xu, Jingxuan, et al.
Published: (2024)
Learning a Particle Dynamics Model with Real-world Videos
by: Kim, Chanho, et al.
Published: (2026)
by: Kim, Chanho, et al.
Published: (2026)
Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
by: Peng, Chuang, et al.
Published: (2025)
by: Peng, Chuang, et al.
Published: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
by: Chen, Boyu, et al.
Published: (2024)
by: Chen, Boyu, et al.
Published: (2024)
CutClaw: Agentic Hours-Long Video Editing via Music Synchronization
by: Zhao, Shifang, et al.
Published: (2026)
by: Zhao, Shifang, et al.
Published: (2026)
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
by: Ge, Yuying, et al.
Published: (2025)
by: Ge, Yuying, et al.
Published: (2025)
Cambrian-P: Pose-Grounded Video Understanding
by: Yang, Jihan, et al.
Published: (2026)
by: Yang, Jihan, et al.
Published: (2026)
Evaluating the Effectiveness of Video Anomaly Detection in the Wild: Online Learning and Inference for Real-world Deployment
by: Yao, Shanle, et al.
Published: (2024)
by: Yao, Shanle, et al.
Published: (2024)
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
by: Li, Longfei, et al.
Published: (2025)
by: Li, Longfei, et al.
Published: (2025)
AvatarMakeup: Realistic Makeup Transfer for 3D Animatable Head Avatars
by: Zhong, Yiming, et al.
Published: (2025)
by: Zhong, Yiming, et al.
Published: (2025)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
by: Xiao, Yuxi, et al.
Published: (2025)
by: Xiao, Yuxi, et al.
Published: (2025)
VideoSAM: Open-World Video Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Similar Items
-
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
by: Ren, Zhongwei, et al.
Published: (2025) -
PixelLM: Pixel Reasoning with Large Multimodal Model
by: Ren, Zhongwei, et al.
Published: (2023) -
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
by: Xing, Ke, et al.
Published: (2025) -
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023) -
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
by: Zhang, Haoji, et al.
Published: (2025)