Subtask-Aware Visual Reward Learning from Segmented Demonstrations
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Changyeon, Heo, Minho, Lee, Doohyun, Shin, Jinwoo, Lee, Honglak, Lim, Joseph J., Lee, Kimin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
RVN-Bench: A Benchmark for Reactive Visual Navigation
by: Lee, Jaewon, et al.
Published: (2026)
by: Lee, Jaewon, et al.
Published: (2026)
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
by: Shin, Ukcheol, et al.
Published: (2023)
by: Shin, Ukcheol, et al.
Published: (2023)
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation
by: Kim, Minwoo, et al.
Published: (2025)
by: Kim, Minwoo, et al.
Published: (2025)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
by: Kim, Seungku, et al.
Published: (2026)
by: Kim, Seungku, et al.
Published: (2026)
Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning
by: Shin, Ukcheol, et al.
Published: (2025)
by: Shin, Ukcheol, et al.
Published: (2025)
DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion Models
by: Chae, Daewon, et al.
Published: (2025)
by: Chae, Daewon, et al.
Published: (2025)
Personalized Reward Modeling for Text-to-Image Generation
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions
by: Lee, Jihyun, et al.
Published: (2026)
by: Lee, Jihyun, et al.
Published: (2026)
A Comparative Study of Machine Learning and Deep Learning for Out-of-Distribution Detection
by: Baek, Jihyeon, et al.
Published: (2026)
by: Baek, Jihyeon, et al.
Published: (2026)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
by: Back, Seunghyeok, et al.
Published: (2025)
by: Back, Seunghyeok, et al.
Published: (2025)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
by: Kim, Jisoo, et al.
Published: (2026)
by: Kim, Jisoo, et al.
Published: (2026)
Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning
by: Lee, Andrew, et al.
Published: (2025)
by: Lee, Andrew, et al.
Published: (2025)
Learning to Control Camera Exposure via Reinforcement Learning
by: Lee, Kyunghyun, et al.
Published: (2024)
by: Lee, Kyunghyun, et al.
Published: (2024)
Spanning Tree Autoregressive Visual Generation
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Trust Region Q Adjoint Matching
by: Dong, Yonghoon, et al.
Published: (2026)
by: Dong, Yonghoon, et al.
Published: (2026)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026)
by: Lee, Seokmin, et al.
Published: (2026)
Mutually-Aware Feature Learning for Few-Shot Object Counting
by: Jeon, Yerim, et al.
Published: (2024)
by: Jeon, Yerim, et al.
Published: (2024)
Efficient Perception, Planning, and Control Algorithm for Vision-Based Automated Vehicles
by: Lee, Der-Hau
Published: (2022)
by: Lee, Der-Hau
Published: (2022)
AiSDF: Structure-aware Neural Signed Distance Fields in Indoor Scenes
by: Jang, Jaehoon, et al.
Published: (2024)
by: Jang, Jaehoon, et al.
Published: (2024)
InstructBooth: Instruction-following Personalized Text-to-Image Generation
by: Chae, Daewon, et al.
Published: (2023)
by: Chae, Daewon, et al.
Published: (2023)
Embodied Uncertainty-Aware Object Segmentation
by: Fang, Xiaolin, et al.
Published: (2024)
by: Fang, Xiaolin, et al.
Published: (2024)
RVT-2: Learning Precise Manipulation from Few Demonstrations
by: Goyal, Ankit, et al.
Published: (2024)
by: Goyal, Ankit, et al.
Published: (2024)
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
by: Kim, Younghyun, et al.
Published: (2025)
by: Kim, Younghyun, et al.
Published: (2025)
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
by: Yang, Yu, et al.
Published: (2025)
by: Yang, Yu, et al.
Published: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
by: Perincherry, Akhil, et al.
Published: (2025)
by: Perincherry, Akhil, et al.
Published: (2025)
Quadrotor Navigation using Reinforcement Learning with Privileged Information
by: Lee, Jonathan, et al.
Published: (2025)
by: Lee, Jonathan, et al.
Published: (2025)
g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations
by: Huang, Wei-Jin, et al.
Published: (2026)
by: Huang, Wei-Jin, et al.
Published: (2026)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025)
by: Zahid, Azizul, et al.
Published: (2025)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Zero-Shot Industrial Anomaly Segmentation with Image-Aware Prompt Generation
by: Park, SoYoung, et al.
Published: (2025)
by: Park, SoYoung, et al.
Published: (2025)
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
by: Chuang, Ian, et al.
Published: (2025)
by: Chuang, Ian, et al.
Published: (2025)
Similar Items
-
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
by: Koo, Myungkyu, et al.
Published: (2025) -
RVN-Bench: A Benchmark for Reactive Visual Navigation
by: Lee, Jaewon, et al.
Published: (2026) -
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024) -
Complementary Random Masking for RGB-Thermal Semantic Segmentation
by: Shin, Ukcheol, et al.
Published: (2023) -
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation
by: Kim, Minwoo, et al.
Published: (2025)