STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Zhonghan, Chai, Wenhao, Wang, Xuan, Ma, Ke, Chen, Kewei, Guo, Dongxu, Ye, Tian, Zhang, Yanting, Wang, Hongwei, Wang, Gaoang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
STEVE: A Step Verification Pipeline for Computer-use Agent Training
von: Lu, Fanbin, et al.
Veröffentlicht: (2025)
von: Lu, Fanbin, et al.
Veröffentlicht: (2025)
DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target
von: Hu, BoCheng, et al.
Veröffentlicht: (2026)
von: Hu, BoCheng, et al.
Veröffentlicht: (2026)
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning
von: Wang, Xuan, et al.
Veröffentlicht: (2023)
von: Wang, Xuan, et al.
Veröffentlicht: (2023)
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
von: Xu, Weili, et al.
Veröffentlicht: (2025)
von: Xu, Weili, et al.
Veröffentlicht: (2025)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
von: Guo, Xuechen, et al.
Veröffentlicht: (2024)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
von: Hao, Shengyu, et al.
Veröffentlicht: (2024)
von: Hao, Shengyu, et al.
Veröffentlicht: (2024)
CityCraft: A Real Crafter for 3D City Generation
von: Deng, Jie, et al.
Veröffentlicht: (2024)
von: Deng, Jie, et al.
Veröffentlicht: (2024)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
von: Tang, Tianci, et al.
Veröffentlicht: (2026)
von: Tang, Tianci, et al.
Veröffentlicht: (2026)
MPM: A Unified 2D-3D Human Pose Representation via Masked Pose Modeling
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2023)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2023)
Blind Inpainting with Object-aware Discrimination for Artificial Marker Removal
von: Guo, Xuechen, et al.
Veröffentlicht: (2023)
von: Guo, Xuechen, et al.
Veröffentlicht: (2023)
Step by Step Network
von: Han, Dongchen, et al.
Veröffentlicht: (2025)
von: Han, Dongchen, et al.
Veröffentlicht: (2025)
See and Think: Embodied Agent in Virtual Environment
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2023)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2023)
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
von: Song, Enxin, et al.
Veröffentlicht: (2023)
von: Song, Enxin, et al.
Veröffentlicht: (2023)
FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
VersaT2I: Improving Text-to-Image Models with Versatile Reward
von: Guo, Jianshu, et al.
Veröffentlicht: (2024)
von: Guo, Jianshu, et al.
Veröffentlicht: (2024)
Hi-LSplat: Hierarchical 3D Language Gaussian Splatting
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
von: Zhang, Lechao, et al.
Veröffentlicht: (2026)
von: Zhang, Lechao, et al.
Veröffentlicht: (2026)
DSG-World: Learning a 3D Gaussian World Model from Dual State Videos
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis
von: Shao, Shitong, et al.
Veröffentlicht: (2025)
von: Shao, Shitong, et al.
Veröffentlicht: (2025)
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
von: Song, Enxin, et al.
Veröffentlicht: (2025)
von: Song, Enxin, et al.
Veröffentlicht: (2025)
RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
von: Zhan, Chenlu, et al.
Veröffentlicht: (2025)
Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
CityGen: Infinite and Controllable City Layout Generation
von: Deng, Jie, et al.
Veröffentlicht: (2023)
von: Deng, Jie, et al.
Veröffentlicht: (2023)
Hand3R: Online 4D Hand-Scene Reconstruction in the Wild
von: Hu, Wendi, et al.
Veröffentlicht: (2026)
von: Hu, Wendi, et al.
Veröffentlicht: (2026)
X-MoGen: Unified Motion Generation across Humans and Animals
von: Wang, Xuan, et al.
Veröffentlicht: (2025)
von: Wang, Xuan, et al.
Veröffentlicht: (2025)
DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
von: Chen, Zheng, et al.
Veröffentlicht: (2025)
PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
von: Wang, Siting, et al.
Veröffentlicht: (2026)
von: Wang, Siting, et al.
Veröffentlicht: (2026)
Disjoint Contrastive Regression Learning for Multi-Sourced Annotations
von: Ruan, Xiaoqian, et al.
Veröffentlicht: (2021)
von: Ruan, Xiaoqian, et al.
Veröffentlicht: (2021)
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
von: Wang, Shuyun, et al.
Veröffentlicht: (2025)
von: Wang, Shuyun, et al.
Veröffentlicht: (2025)
A Step to Decouple Optimization in 3DGS
von: Ding, Renjie, et al.
Veröffentlicht: (2026)
von: Ding, Renjie, et al.
Veröffentlicht: (2026)
MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant
von: Zhan, Chenlu, et al.
Veröffentlicht: (2024)
von: Zhan, Chenlu, et al.
Veröffentlicht: (2024)
Learning Human Skill Generators at Key-Step Levels
von: Wu, Yilu, et al.
Veröffentlicht: (2025)
von: Wu, Yilu, et al.
Veröffentlicht: (2025)
Real-time One-Step Diffusion-based Expressive Portrait Videos Generation
von: Guo, Hanzhong, et al.
Veröffentlicht: (2024)
von: Guo, Hanzhong, et al.
Veröffentlicht: (2024)
Concept Unlearning by Modeling Key Steps of Diffusion Process
von: Zhang, Chaoshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Chaoshuo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024) -
Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024) -
STEVE: A Step Verification Pipeline for Computer-use Agent Training
von: Lu, Fanbin, et al.
Veröffentlicht: (2025) -
DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target
von: Hu, BoCheng, et al.
Veröffentlicht: (2026) -
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning
von: Wang, Xuan, et al.
Veröffentlicht: (2023)