Latent Action Pretraining from Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Seonghyeon, Jang, Joel, Jeon, Byeongguk, Joo, Sejune, Yang, Jianwei, Peng, Baolin, Mandlekar, Ajay, Tan, Reuben, Chao, Yu-Wei, Lin, Bill Yuchen, Liden, Lars, Lee, Kimin, Gao, Jianfeng, Zettlemoyer, Luke, Fox, Dieter, Seo, Minjoon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
by: Doo, JaeHyeok, et al.
Published: (2026)
by: Doo, JaeHyeok, et al.
Published: (2026)
Prime the search: Using large language models for guiding geometric task and motion planning by warm-starting tree search
by: Lee, Dongryung, et al.
Published: (2025)
by: Lee, Dongryung, et al.
Published: (2025)
How Well Do Large Language Models Truly Ground?
by: Lee, Hyunji, et al.
Published: (2023)
by: Lee, Hyunji, et al.
Published: (2023)
Semiparametric Token-Sequence Co-Supervision
by: Lee, Hyunji, et al.
Published: (2024)
by: Lee, Hyunji, et al.
Published: (2024)
Magma: A Foundation Model for Multimodal AI Agents
by: Yang, Jianwei, et al.
Published: (2025)
by: Yang, Jianwei, et al.
Published: (2025)
SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment
by: Garrett, Caelan, et al.
Published: (2024)
by: Garrett, Caelan, et al.
Published: (2024)
AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback
by: Tupini, Andrea, et al.
Published: (2026)
by: Tupini, Andrea, et al.
Published: (2026)
Latent Reasoning via Sentence Embedding Prediction
by: Hwang, Hyeonbin, et al.
Published: (2025)
by: Hwang, Hyeonbin, et al.
Published: (2025)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
by: Chang, Hoyeon, et al.
Published: (2024)
by: Chang, Hoyeon, et al.
Published: (2024)
IntervenGen: Interventional Data Generation for Robust and Data-Efficient Robot Imitation Learning
by: Hoque, Ryan, et al.
Published: (2024)
by: Hoque, Ryan, et al.
Published: (2024)
SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation
by: Zhou, Zihan, et al.
Published: (2024)
by: Zhou, Zihan, et al.
Published: (2024)
Improving Probability-based Prompt Selection Through Unified Evaluation and Analysis
by: Yang, Sohee, et al.
Published: (2023)
by: Yang, Sohee, et al.
Published: (2023)
Point Bridge: 3D Representations for Cross Domain Policy Learning
by: Haldar, Siddhant, et al.
Published: (2026)
by: Haldar, Siddhant, et al.
Published: (2026)
DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation
by: Mandi, Zhao, et al.
Published: (2025)
by: Mandi, Zhao, et al.
Published: (2025)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
by: Hwang, Hyeonbin, et al.
Published: (2024)
by: Hwang, Hyeonbin, et al.
Published: (2024)
INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
by: Oh, Hanseok, et al.
Published: (2024)
by: Oh, Hanseok, et al.
Published: (2024)
Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
by: Kim, Geewook, et al.
Published: (2025)
by: Kim, Geewook, et al.
Published: (2025)
State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
by: Kim, Geewook, et al.
Published: (2025)
by: Kim, Geewook, et al.
Published: (2025)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
by: Kim, Geewook, et al.
Published: (2024)
by: Kim, Geewook, et al.
Published: (2024)
Rethinking the Role of Proxy Rewards in Language Model Alignment
by: Kim, Sungdong, et al.
Published: (2024)
by: Kim, Sungdong, et al.
Published: (2024)
Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition
by: Kim, Jiyeon, et al.
Published: (2024)
by: Kim, Jiyeon, et al.
Published: (2024)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
by: Zhou, Zihan, et al.
Published: (2025)
by: Zhou, Zihan, et al.
Published: (2025)
NOD-TAMP: Generalizable Long-Horizon Planning with Neural Object Descriptors
by: Cheng, Shuo, et al.
Published: (2023)
by: Cheng, Shuo, et al.
Published: (2023)
DSAI: Unbiased and Interpretable Latent Feature Extraction for Data-Centric AI
by: Cho, Hyowon, et al.
Published: (2024)
by: Cho, Hyowon, et al.
Published: (2024)
Trabajo forestal en régimen de subcontratación en Suecia
by: Ewa Lidén (Author)
Published: (1997)
by: Ewa Lidén (Author)
Published: (1997)
Le travail en sous-traitance dans la foresterie en Suède
by: Ewa Lidén (Author)
Published: (1997)
by: Ewa Lidén (Author)
Published: (1997)
Contract labour in forestry in Sweden
by: Ewa Lidén (Author)
Published: (1997)
by: Ewa Lidén (Author)
Published: (1997)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
by: Shi, Weijia, et al.
Published: (2024)
by: Shi, Weijia, et al.
Published: (2024)
Ask Optimal Questions: Aligning Large Language Models with Retriever's Preference in Conversation
by: Yoon, Chanwoong, et al.
Published: (2024)
by: Yoon, Chanwoong, et al.
Published: (2024)
TSLM: Tree-Structured Language Modeling for Divergent Thinking
by: Kim, Doyoung, et al.
Published: (2026)
by: Kim, Doyoung, et al.
Published: (2026)
Detecting Pretraining Data from Large Language Models
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
From What to Respond to When to Respond: Timely Response Generation for Open-domain Dialogue Agents
by: Jang, Seongbo, et al.
Published: (2025)
by: Jang, Seongbo, et al.
Published: (2025)
KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
by: Jang, Seongbo, et al.
Published: (2024)
by: Jang, Seongbo, et al.
Published: (2024)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
by: Ye, Seonghyeon, et al.
Published: (2023)
by: Ye, Seonghyeon, et al.
Published: (2023)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024)
by: Yoon, Dongkeun, et al.
Published: (2024)
Exploring the Practicality of Generative Retrieval on Dynamic Corpora
by: Kim, Chaeeun, et al.
Published: (2023)
by: Kim, Chaeeun, et al.
Published: (2023)
Chapter 9 Mopping Up, Keeping Down, and Propping Up
by: Lidén, Kristoffer, et al.
Published: (2024)
by: Lidén, Kristoffer, et al.
Published: (2024)
Signatures Meet Dynamic Programming: Generalizing Bellman Equations for Trajectory Following
by: Ohnishi, Motoya, et al.
Published: (2023)
by: Ohnishi, Motoya, et al.
Published: (2023)
Similar Items
-
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
by: Doo, JaeHyeok, et al.
Published: (2026) -
Prime the search: Using large language models for guiding geometric task and motion planning by warm-starting tree search
by: Lee, Dongryung, et al.
Published: (2025) -
How Well Do Large Language Models Truly Ground?
by: Lee, Hyunji, et al.
Published: (2023) -
Semiparametric Token-Sequence Co-Supervision
by: Lee, Hyunji, et al.
Published: (2024) -
Magma: A Foundation Model for Multimodal AI Agents
by: Yang, Jianwei, et al.
Published: (2025)