VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ren, Zhongwei, Wei, Yunchao, Guo, Xun, Zhao, Yao, Kang, Bingyi, Feng, Jiashi, Jin, Xiaojie |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
par: Ren, Zhongwei, et autres
Publié: (2026)
par: Ren, Zhongwei, et autres
Publié: (2026)
PixelLM: Pixel Reasoning with Large Multimodal Model
par: Ren, Zhongwei, et autres
Publié: (2023)
par: Ren, Zhongwei, et autres
Publié: (2023)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
par: Yang, Lihe, et autres
Publié: (2024)
par: Yang, Lihe, et autres
Publié: (2024)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
par: Chen, Sili, et autres
Publié: (2025)
par: Chen, Sili, et autres
Publié: (2025)
PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
par: Yin, Yuyang, et autres
Publié: (2025)
par: Yin, Yuyang, et autres
Publié: (2025)
How Far is Video Generation from World Model: A Physical Law Perspective
par: Kang, Bingyi, et autres
Publié: (2024)
par: Kang, Bingyi, et autres
Publié: (2024)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
par: Jin, Xiaojie, et autres
Publié: (2023)
par: Jin, Xiaojie, et autres
Publié: (2023)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
par: Wang, Yuqing, et autres
Publié: (2024)
par: Wang, Yuqing, et autres
Publié: (2024)
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
par: Xing, Ke, et autres
Publié: (2025)
par: Xing, Ke, et autres
Publié: (2025)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
par: Zhang, Haoji, et autres
Publié: (2025)
par: Zhang, Haoji, et autres
Publié: (2025)
Hierarchical Memory for Long Video QA
par: Wang, Yiqin, et autres
Publié: (2024)
par: Wang, Yiqin, et autres
Publié: (2024)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
par: Ma, Fan, et autres
Publié: (2023)
par: Ma, Fan, et autres
Publié: (2023)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
par: Liu, Xinhang, et autres
Publié: (2025)
par: Liu, Xinhang, et autres
Publié: (2025)
A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
par: Lin, Dongheng, et autres
Publié: (2025)
par: Lin, Dongheng, et autres
Publié: (2025)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
par: Zhang, Haoji, et autres
Publié: (2024)
par: Zhang, Haoji, et autres
Publié: (2024)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
par: Yu, Xiao, et autres
Publié: (2025)
par: Yu, Xiao, et autres
Publié: (2025)
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
par: Li, Zhenyu, et autres
Publié: (2025)
par: Li, Zhenyu, et autres
Publié: (2025)
Classification Done Right for Vision-Language Pre-Training
par: Huang, Zilong, et autres
Publié: (2024)
par: Huang, Zilong, et autres
Publié: (2024)
Image Understanding Makes for A Good Tokenizer for Image Generation
par: Wang, Luting, et autres
Publié: (2024)
par: Wang, Luting, et autres
Publié: (2024)
Depth Anything V2
par: Yang, Lihe, et autres
Publié: (2024)
par: Yang, Lihe, et autres
Publié: (2024)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
par: Xu, Xiaojie, et autres
Publié: (2026)
par: Xu, Xiaojie, et autres
Publié: (2026)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
par: Xu, Lin, et autres
Publié: (2024)
par: Xu, Lin, et autres
Publié: (2024)
VRAG: Learning World Models for Interactive Video Generation
par: Chen, Taiye, et autres
Publié: (2025)
par: Chen, Taiye, et autres
Publié: (2025)
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
par: Sun, Yihong, et autres
Publié: (2024)
par: Sun, Yihong, et autres
Publié: (2024)
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
par: Bao, Peijun, et autres
Publié: (2024)
par: Bao, Peijun, et autres
Publié: (2024)
CutClaw: Agentic Hours-Long Video Editing via Music Synchronization
par: Zhao, Shifang, et autres
Publié: (2026)
par: Zhao, Shifang, et autres
Publié: (2026)
Cambrian-P: Pose-Grounded Video Understanding
par: Yang, Jihan, et autres
Publié: (2026)
par: Yang, Jihan, et autres
Publié: (2026)
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
par: Li, Longfei, et autres
Publié: (2025)
par: Li, Longfei, et autres
Publié: (2025)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
par: Majumder, Sagnik, et autres
Publié: (2024)
par: Majumder, Sagnik, et autres
Publié: (2024)
VideoSAM: Open-World Video Segmentation
par: Guo, Pinxue, et autres
Publié: (2024)
par: Guo, Pinxue, et autres
Publié: (2024)
Practical Video Object Detection via Feature Selection and Aggregation
par: Shi, Yuheng, et autres
Publié: (2024)
par: Shi, Yuheng, et autres
Publié: (2024)
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
par: Chen, Yubin, et autres
Publié: (2025)
par: Chen, Yubin, et autres
Publié: (2025)
Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval
par: Dong, Jianfeng, et autres
Publié: (2025)
par: Dong, Jianfeng, et autres
Publié: (2025)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
par: Zhou, Yupeng, et autres
Publié: (2024)
par: Zhou, Yupeng, et autres
Publié: (2024)
PSVMA+: Exploring Multi-granularity Semantic-visual Adaption for Generalized Zero-shot Learning
par: Liu, Man, et autres
Publié: (2024)
par: Liu, Man, et autres
Publié: (2024)
Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
par: Peng, Chuang, et autres
Publié: (2025)
par: Peng, Chuang, et autres
Publié: (2025)
Collaborative Feature-Logits Contrastive Learning for Open-Set Semi-Supervised Object Detection
par: Zhong, Xinhao, et autres
Publié: (2024)
par: Zhong, Xinhao, et autres
Publié: (2024)
Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
par: Sun, Keqiang, et autres
Publié: (2023)
par: Sun, Keqiang, et autres
Publié: (2023)
GV-VAD : Exploring Video Generation for Weakly-Supervised Video Anomaly Detection
par: Cai, Suhang, et autres
Publié: (2025)
par: Cai, Suhang, et autres
Publié: (2025)
Scalable Video Object Segmentation with Identification Mechanism
par: Yang, Zongxin, et autres
Publié: (2022)
par: Yang, Zongxin, et autres
Publié: (2022)
Documents similaires
-
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
par: Ren, Zhongwei, et autres
Publié: (2026) -
PixelLM: Pixel Reasoning with Large Multimodal Model
par: Ren, Zhongwei, et autres
Publié: (2023) -
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
par: Yang, Lihe, et autres
Publié: (2024) -
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
par: Chen, Sili, et autres
Publié: (2025) -
PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
par: Yin, Yuyang, et autres
Publié: (2025)