Visual Representation Learning with Stochastic Frame Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Jang, Huiwon, Kim, Dongyoung, Kim, Junsu, Shin, Jinwoo, Abbeel, Pieter, Seo, Younggyo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
by: Kim, Dongyoung, et al.
Published: (2025)
by: Kim, Dongyoung, et al.
Published: (2025)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
by: Won, John, et al.
Published: (2025)
by: Won, John, et al.
Published: (2025)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
by: Kim, Seungku, et al.
Published: (2026)
by: Kim, Seungku, et al.
Published: (2026)
Continuous Control with Coarse-to-fine Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
by: Vosylius, Vitalis, et al.
Published: (2024)
by: Vosylius, Vitalis, et al.
Published: (2024)
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
by: Jang, Huiwon, et al.
Published: (2025)
by: Jang, Huiwon, et al.
Published: (2025)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
RAPID: Robust and Agile Planner Using Inverse Reinforcement Learning for Vision-Based Drone Navigation
by: Kim, Minwoo, et al.
Published: (2025)
by: Kim, Minwoo, et al.
Published: (2025)
Verifier-free Test-Time Sampling for Vision Language Action Models
by: Jang, Suhyeok, et al.
Published: (2025)
by: Jang, Suhyeok, et al.
Published: (2025)
BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark
by: Chernyadev, Nikita, et al.
Published: (2024)
by: Chernyadev, Nikita, et al.
Published: (2024)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
Twisting Lids Off with Two Hands
by: Lin, Toru, et al.
Published: (2024)
by: Lin, Toru, et al.
Published: (2024)
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
by: Majumdar, Arjun, et al.
Published: (2023)
by: Majumdar, Arjun, et al.
Published: (2023)
How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference
by: Lin, Toru, et al.
Published: (2026)
by: Lin, Toru, et al.
Published: (2026)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
by: Mishra, Nikhil, et al.
Published: (2024)
by: Mishra, Nikhil, et al.
Published: (2024)
Adversarial Robustification via Text-to-Image Diffusion Models
by: Choi, Daewon, et al.
Published: (2024)
by: Choi, Daewon, et al.
Published: (2024)
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
by: Tsagkas, Nikolaos, et al.
Published: (2025)
by: Tsagkas, Nikolaos, et al.
Published: (2025)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
Confidence-aware Denoised Fine-tuning of Off-the-shelf Models for Certified Robustness
by: Jang, Suhyeok, et al.
Published: (2024)
by: Jang, Suhyeok, et al.
Published: (2024)
Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions
by: Kim, Kyungmin, et al.
Published: (2024)
by: Kim, Kyungmin, et al.
Published: (2024)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
by: Das, Richeek, et al.
Published: (2025)
by: Das, Richeek, et al.
Published: (2025)
Learning to Visually Connect Actions and their Effects
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?
by: Mu, Tongzhou, et al.
Published: (2024)
by: Mu, Tongzhou, et al.
Published: (2024)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
by: Jang, Jinhyeok, et al.
Published: (2025)
by: Jang, Jinhyeok, et al.
Published: (2025)
Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations
by: Kim, Minung, et al.
Published: (2025)
by: Kim, Minung, et al.
Published: (2025)
AnyThermal: Towards Learning Universal Representations for Thermal Perception
by: Maheshwari, Parv, et al.
Published: (2026)
by: Maheshwari, Parv, et al.
Published: (2026)
Learned Visual Navigation for Under-Canopy Agricultural Robots
by: Sivakumar, Arun Narenthiran, et al.
Published: (2021)
by: Sivakumar, Arun Narenthiran, et al.
Published: (2021)
ReCoRe: Regularized Contrastive Representation Learning of World Model
by: Poudel, Rudra P. K., et al.
Published: (2023)
by: Poudel, Rudra P. K., et al.
Published: (2023)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
by: Zhou, Siyuan, et al.
Published: (2026)
by: Zhou, Siyuan, et al.
Published: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
by: Kim, Dongyoung, et al.
Published: (2026)
by: Kim, Dongyoung, et al.
Published: (2026)
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
by: Brown, Alexandre, et al.
Published: (2025)
by: Brown, Alexandre, et al.
Published: (2025)
Similar Items
-
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
by: Jang, Huiwon, et al.
Published: (2024) -
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023) -
Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
by: Kim, Dongyoung, et al.
Published: (2025) -
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024) -
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
by: Won, John, et al.
Published: (2025)