Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement
Fuente:
arXiv
Salvato in:
| Autori principali: | Kee, Hogun, Oh, Wooseok, Kang, Minjae, Ahn, Hyemin, Oh, Songhwai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RVN-Bench: A Benchmark for Reactive Visual Navigation
di: Lee, Jaewon, et al.
Pubblicazione: (2026)
di: Lee, Jaewon, et al.
Pubblicazione: (2026)
May the Dance be with You: Dance Generation Framework for Non-Humanoids
di: Ahn, Hyemin
Pubblicazione: (2024)
di: Ahn, Hyemin
Pubblicazione: (2024)
Knolling Bot: Teaching Robots the Human Notion of Tidiness
di: Hu, Yuhang, et al.
Pubblicazione: (2023)
di: Hu, Yuhang, et al.
Pubblicazione: (2023)
Modality-Augmented Fine-Tuning of Foundation Robot Policies for Cross-Embodiment Manipulation on GR1 and G1
di: Park, Junsung, et al.
Pubblicazione: (2025)
di: Park, Junsung, et al.
Pubblicazione: (2025)
TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning
di: Wu, Jimmy, et al.
Pubblicazione: (2024)
di: Wu, Jimmy, et al.
Pubblicazione: (2024)
Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning
di: Shin, Ukcheol, et al.
Pubblicazione: (2025)
di: Shin, Ukcheol, et al.
Pubblicazione: (2025)
RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
di: Sun, Xiaoquan, et al.
Pubblicazione: (2025)
di: Sun, Xiaoquan, et al.
Pubblicazione: (2025)
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
di: Kim, Seo Hyun, et al.
Pubblicazione: (2026)
di: Kim, Seo Hyun, et al.
Pubblicazione: (2026)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
di: Shin, Ukcheol, et al.
Pubblicazione: (2023)
di: Shin, Ukcheol, et al.
Pubblicazione: (2023)
FIReStereo: Forest InfraRed Stereo Dataset for UAS Depth Perception in Visually Degraded Environments
di: Dhrafani, Devansh, et al.
Pubblicazione: (2024)
di: Dhrafani, Devansh, et al.
Pubblicazione: (2024)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
di: Park, Yohan, et al.
Pubblicazione: (2025)
di: Park, Yohan, et al.
Pubblicazione: (2025)
FloNa: Floor Plan Guided Embodied Visual Navigation
di: Li, Jiaxin, et al.
Pubblicazione: (2024)
di: Li, Jiaxin, et al.
Pubblicazione: (2024)
Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning
di: Lee, Andrew, et al.
Pubblicazione: (2025)
di: Lee, Andrew, et al.
Pubblicazione: (2025)
Stabilizing Reinforcement Learning in Differentiable Multiphysics Simulation
di: Xing, Eliot, et al.
Pubblicazione: (2024)
di: Xing, Eliot, et al.
Pubblicazione: (2024)
Robotic Visual Instruction
di: Li, Yanbang, et al.
Pubblicazione: (2025)
di: Li, Yanbang, et al.
Pubblicazione: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
di: Zhang, Yi, et al.
Pubblicazione: (2025)
di: Zhang, Yi, et al.
Pubblicazione: (2025)
GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation
di: Wu, Ruihai, et al.
Pubblicazione: (2025)
di: Wu, Ruihai, et al.
Pubblicazione: (2025)
Prior Availability in Industrial Visual Sim-to-Real: A Review of CAD-Guided and CAD-Unavailable Regimes
di: Tao, Chenxi, et al.
Pubblicazione: (2026)
di: Tao, Chenxi, et al.
Pubblicazione: (2026)
LINGO-Space: Language-Conditioned Incremental Grounding for Space
di: Kim, Dohyun, et al.
Pubblicazione: (2024)
di: Kim, Dohyun, et al.
Pubblicazione: (2024)
Instruction-Guided Visual Masking
di: Zheng, Jinliang, et al.
Pubblicazione: (2024)
di: Zheng, Jinliang, et al.
Pubblicazione: (2024)
Accelerating Transformer-Based Monocular SLAM via Geometric Utility Scoring
di: Xiong, Xinmiao, et al.
Pubblicazione: (2026)
di: Xiong, Xinmiao, et al.
Pubblicazione: (2026)
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2025)
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2025)
ManiSkill-HAB: A Benchmark for Low-Level Manipulation in Home Rearrangement Tasks
di: Shukla, Arth, et al.
Pubblicazione: (2024)
di: Shukla, Arth, et al.
Pubblicazione: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
di: Kang, Wonjun, et al.
Pubblicazione: (2025)
di: Kang, Wonjun, et al.
Pubblicazione: (2025)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
di: Kim, Sumin, et al.
Pubblicazione: (2026)
di: Kim, Sumin, et al.
Pubblicazione: (2026)
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
di: Zhong, Fangwei, et al.
Pubblicazione: (2024)
di: Zhong, Fangwei, et al.
Pubblicazione: (2024)
On Extending Semantic Abstraction for Efficient Search of Hidden Objects
di: Pais, Tasha, et al.
Pubblicazione: (2025)
di: Pais, Tasha, et al.
Pubblicazione: (2025)
DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
di: Azhari, Maulana Bisyir, et al.
Pubblicazione: (2025)
di: Azhari, Maulana Bisyir, et al.
Pubblicazione: (2025)
Safe CoR: A Dual-Expert Approach to Integrating Imitation Learning and Safe Reinforcement Learning Using Constraint Rewards
di: Kwon, Hyeokjin, et al.
Pubblicazione: (2024)
di: Kwon, Hyeokjin, et al.
Pubblicazione: (2024)
Semantic Environment Atlas for Object-Goal Navigation
di: Kim, Nuri, et al.
Pubblicazione: (2024)
di: Kim, Nuri, et al.
Pubblicazione: (2024)
DVGT: Driving Visual Geometry Transformer
di: Zuo, Sicheng, et al.
Pubblicazione: (2025)
di: Zuo, Sicheng, et al.
Pubblicazione: (2025)
Object-Centric World Model for Language-Guided Manipulation
di: Jeong, Youngjoon, et al.
Pubblicazione: (2025)
di: Jeong, Youngjoon, et al.
Pubblicazione: (2025)
Autonomous Vision-Guided Resection of Central Airway Obstruction
di: Smith, M. E., et al.
Pubblicazione: (2025)
di: Smith, M. E., et al.
Pubblicazione: (2025)
A Unified Masked Autoencoder with Patchified Skeletons for Motion Synthesis
di: Mascaro, Esteve Valls, et al.
Pubblicazione: (2023)
di: Mascaro, Esteve Valls, et al.
Pubblicazione: (2023)
Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search
di: Paramonov, Kirill, et al.
Pubblicazione: (2024)
di: Paramonov, Kirill, et al.
Pubblicazione: (2024)
Visual IRL for Human-Like Robotic Manipulation
di: Asali, Ehsan, et al.
Pubblicazione: (2024)
di: Asali, Ehsan, et al.
Pubblicazione: (2024)
Visual SLAMMOT Considering Multiple Motion Models
di: Tian, Peilin, et al.
Pubblicazione: (2024)
di: Tian, Peilin, et al.
Pubblicazione: (2024)
Language-Conditioned World Modeling for Visual Navigation
di: Dong, Yifei, et al.
Pubblicazione: (2026)
di: Dong, Yifei, et al.
Pubblicazione: (2026)
From Scene to Object: Text-Guided Dual-Gaze Prediction
di: Ke, Zehong, et al.
Pubblicazione: (2026)
di: Ke, Zehong, et al.
Pubblicazione: (2026)
ContactHandover: Contact-Guided Robot-to-Human Object Handover
di: Wang, Zixi, et al.
Pubblicazione: (2024)
di: Wang, Zixi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
RVN-Bench: A Benchmark for Reactive Visual Navigation
di: Lee, Jaewon, et al.
Pubblicazione: (2026) -
May the Dance be with You: Dance Generation Framework for Non-Humanoids
di: Ahn, Hyemin
Pubblicazione: (2024) -
Knolling Bot: Teaching Robots the Human Notion of Tidiness
di: Hu, Yuhang, et al.
Pubblicazione: (2023) -
Modality-Augmented Fine-Tuning of Foundation Robot Policies for Cross-Embodiment Manipulation on GR1 and G1
di: Park, Junsung, et al.
Pubblicazione: (2025) -
TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning
di: Wu, Jimmy, et al.
Pubblicazione: (2024)