Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhai, Yuexiang, Bai, Hao, Lin, Zipeng, Pan, Jiayi, Tong, Shengbang, Zhou, Yifei, Suhr, Alane, Xie, Saining, LeCun, Yann, Ma, Yi, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
by: Bai, Hao, et al.
Published: (2024)
by: Bai, Hao, et al.
Published: (2024)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Scaling Language-Free Visual Representation Learning
by: Fan, David, et al.
Published: (2025)
by: Fan, David, et al.
Published: (2025)
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Autonomous Evaluation and Refinement of Digital Agents
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
A hierarchical loss and its problems when classifying non-hierarchically
by: Wu, Cinna, et al.
Published: (2017)
by: Wu, Cinna, et al.
Published: (2017)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Grounding Language in Multi-Perspective Referential Communication
by: Tang, Zineng, et al.
Published: (2024)
by: Tang, Zineng, et al.
Published: (2024)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
by: Sun, Shangwen, et al.
Published: (2026)
by: Sun, Shangwen, et al.
Published: (2026)
Diffusion Transformers with Representation Autoencoders
by: Zheng, Boyang, et al.
Published: (2025)
by: Zheng, Boyang, et al.
Published: (2025)
Evaluating Model Perception of Color Illusions in Photorealistic Scenes
by: Mao, Lingjun, et al.
Published: (2024)
by: Mao, Lingjun, et al.
Published: (2024)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
by: Yun, Zeyu, et al.
Published: (2023)
by: Yun, Zeyu, et al.
Published: (2023)
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Transformers without Normalization
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
Learning Adaptive Parallel Reasoning with Language Models
by: Pan, Jiayi, et al.
Published: (2025)
by: Pan, Jiayi, et al.
Published: (2025)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)
by: Skean, Oscar, et al.
Published: (2024)
Cambrian-S: Towards Spatial Supersensing in Video
by: Yang, Shusheng, et al.
Published: (2025)
by: Yang, Shusheng, et al.
Published: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
by: Denton, Remi, et al.
Published: (2014)
by: Denton, Remi, et al.
Published: (2014)
Using Language Models to Disambiguate Lexical Choices in Translation
by: Barua, Josh, et al.
Published: (2024)
by: Barua, Josh, et al.
Published: (2024)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Fast and Exact Enumeration of Deep Networks Partitions Regions
by: Balestriero, Randall, et al.
Published: (2024)
by: Balestriero, Randall, et al.
Published: (2024)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
by: Dawid, Anna, et al.
Published: (2023)
by: Dawid, Anna, et al.
Published: (2023)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
by: Sclar, Melanie, et al.
Published: (2023)
by: Sclar, Melanie, et al.
Published: (2023)
Long Chain-of-Thought Reasoning Across Languages
by: Barua, Josh, et al.
Published: (2025)
by: Barua, Josh, et al.
Published: (2025)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
PooDLe: Pooled and dense self-supervised learning from naturalistic videos
by: Wang, Alex N., et al.
Published: (2024)
by: Wang, Alex N., et al.
Published: (2024)
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023)
by: Yu, Yaodong, et al.
Published: (2023)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
by: Zhou, Yifei, et al.
Published: (2024)
by: Zhou, Yifei, et al.
Published: (2024)
Variance-Covariance Regularization Improves Representation Learning
by: Zhu, Jiachen, et al.
Published: (2023)
by: Zhu, Jiachen, et al.
Published: (2023)
Similar Items
-
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
by: Tong, Shengbang, et al.
Published: (2024) -
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
by: Tong, Shengbang, et al.
Published: (2024) -
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025) -
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026) -
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
by: Bai, Hao, et al.
Published: (2024)