Open-World Skill Discovery from Unsegmented Demonstrations
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Jingwen, Wang, Zihao, Cai, Shaofei, Liu, Anji, Liang, Yitao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoopNav: Benchmarking Spatial Consistency in World Models
by: Lian, Kewei, et al.
Published: (2025)
by: Lian, Kewei, et al.
Published: (2025)
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse
by: Li, Muyao, et al.
Published: (2025)
by: Li, Muyao, et al.
Published: (2025)
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023)
by: Zheng, Xinyue, et al.
Published: (2023)
Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
RenderWorld: World Model with Self-Supervised 3D Label
by: Yan, Ziyang, et al.
Published: (2024)
by: Yan, Ziyang, et al.
Published: (2024)
OpenSDI: Spotting Diffusion-Generated Images in the Open World
by: Wang, Yabin, et al.
Published: (2025)
by: Wang, Yabin, et al.
Published: (2025)
RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
by: Yang, Jihan, et al.
Published: (2023)
by: Yang, Jihan, et al.
Published: (2023)
Smart Help: Strategic Opponent Modeling for Proactive and Adaptive Robot Assistance in Households
by: Cao, Zhihao, et al.
Published: (2024)
by: Cao, Zhihao, et al.
Published: (2024)
Unsegment Anything by Simulating Deformation
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks
by: Deng, Yueci, et al.
Published: (2026)
by: Deng, Yueci, et al.
Published: (2026)
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
by: Lin, Wang, et al.
Published: (2026)
by: Lin, Wang, et al.
Published: (2026)
Decoding Decision Reasoning: A Counterfactual-Powered Model for Knowledge Discovery
by: Fang, Yingying, et al.
Published: (2024)
by: Fang, Yingying, et al.
Published: (2024)
OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation
by: Li, Jin, et al.
Published: (2026)
by: Li, Jin, et al.
Published: (2026)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment
by: Garrett, Caelan, et al.
Published: (2024)
by: Garrett, Caelan, et al.
Published: (2024)
Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar Creation
by: Chen, Xiyi, et al.
Published: (2024)
by: Chen, Xiyi, et al.
Published: (2024)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
Skill-Conditioned Visual Geolocation for Vision-Language Models
by: Yang, Chenjie, et al.
Published: (2026)
by: Yang, Chenjie, et al.
Published: (2026)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
by: Liu, Zishan, et al.
Published: (2026)
by: Liu, Zishan, et al.
Published: (2026)
Plug-and-Play Context Feature Reuse for Efficient Masked Generation
by: Liu, Xuejie, et al.
Published: (2025)
by: Liu, Xuejie, et al.
Published: (2025)
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
by: Yang, Deshun, et al.
Published: (2024)
by: Yang, Deshun, et al.
Published: (2024)
CLoG: Benchmarking Continual Learning of Image Generation Models
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision
by: Wang, Junjie, et al.
Published: (2024)
by: Wang, Junjie, et al.
Published: (2024)
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
by: Guo, Junliang, et al.
Published: (2025)
by: Guo, Junliang, et al.
Published: (2025)
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
by: Yang, Yu, et al.
Published: (2025)
by: Yang, Yu, et al.
Published: (2025)
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Bird-SR: Bidirectional Reward-Guided Diffusion for Real-World Image Super-Resolution
by: Fan, Zihao, et al.
Published: (2026)
by: Fan, Zihao, et al.
Published: (2026)
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
From Open Vocabulary to Open World: Teaching Vision Language Models to Detect Novel Objects
by: Li, Zizhao, et al.
Published: (2024)
by: Li, Zizhao, et al.
Published: (2024)
A World Model of Radiologist Reading for Medical Image Representation Learning
by: Li, Yiwei, et al.
Published: (2026)
by: Li, Yiwei, et al.
Published: (2026)
Open-World Motion Forecasting
by: Schischka, Nicolas, et al.
Published: (2026)
by: Schischka, Nicolas, et al.
Published: (2026)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
by: Qin, Jiahao, et al.
Published: (2023)
by: Qin, Jiahao, et al.
Published: (2023)
Unlocking Transfer Learning for Open-World Few-Shot Recognition
by: Kim, Byeonggeun, et al.
Published: (2024)
by: Kim, Byeonggeun, et al.
Published: (2024)
EffOWT: Transfer Visual Language Models to Open-World Tracking Efficiently and Effectively
by: Wang, Bingyang, et al.
Published: (2025)
by: Wang, Bingyang, et al.
Published: (2025)
SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
by: Shi, Yukai, et al.
Published: (2025)
by: Shi, Yukai, et al.
Published: (2025)
Exclusive Style Removal for Cross Domain Novel Class Discovery
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
by: Zheng, Chaoda, et al.
Published: (2026)
by: Zheng, Chaoda, et al.
Published: (2026)
Pose-Star: Anatomy-Aware Editing for Open-World Fashion Images
by: Dong, Yuran, et al.
Published: (2025)
by: Dong, Yuran, et al.
Published: (2025)
Similar Items
-
LoopNav: Benchmarking Spatial Consistency in World Models
by: Lian, Kewei, et al.
Published: (2025) -
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024) -
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
by: Cai, Shaofei, et al.
Published: (2025) -
JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse
by: Li, Muyao, et al.
Published: (2025) -
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023)