Beyond Language Modeling: An Exploration of Multimodal Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tong, Shengbang, Fan, David, Nguyen, John, Brown, Ellis, Zhou, Gaoyue, Qian, Shengyi, Zheng, Boyang, Vallaeys, Théophane, Han, Junlin, Fergus, Rob, Murray, Naila, Ghazvininejad, Marjan, Lewis, Mike, Ballas, Nicolas, Bar, Amir, Rabbat, Michael, Verbeek, Jakob, Zettlemoyer, Luke, Sinha, Koustuv, LeCun, Yann, Xie, Saining |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Language-Free Visual Representation Learning
von: Fan, David, et al.
Veröffentlicht: (2025)
von: Fan, David, et al.
Veröffentlicht: (2025)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
von: Tong, Shengbang, et al.
Veröffentlicht: (2026)
von: Tong, Shengbang, et al.
Veröffentlicht: (2026)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2026)
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2026)
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2025)
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2025)
Learning Latent Action World Models In The Wild
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
Navigation World Models
von: Bar, Amir, et al.
Veröffentlicht: (2024)
von: Bar, Amir, et al.
Veröffentlicht: (2024)
Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2025)
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2025)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2024)
von: Vallaeys, Théophane, et al.
Veröffentlicht: (2024)
Parallel Stochastic Gradient-Based Planning for World Models
von: Psenka, Michael, et al.
Veröffentlicht: (2026)
von: Psenka, Michael, et al.
Veröffentlicht: (2026)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
von: Yasunaga, Michihiro, et al.
Veröffentlicht: (2025)
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
von: Zhou, Gaoyue, et al.
Veröffentlicht: (2024)
von: Zhou, Gaoyue, et al.
Veröffentlicht: (2024)
VUGEN: Visual Understanding priors for GENeration
von: Chen, Xiangyi, et al.
Veröffentlicht: (2025)
von: Chen, Xiangyi, et al.
Veröffentlicht: (2025)
Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
von: Denton, Remi, et al.
Veröffentlicht: (2014)
von: Denton, Remi, et al.
Veröffentlicht: (2014)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
von: Terver, Basile, et al.
Veröffentlicht: (2026)
von: Terver, Basile, et al.
Veröffentlicht: (2026)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2025)
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2025)
Revisiting Feature Prediction for Learning Visual Representations from Video
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
Cambrian-S: Towards Spatial Supersensing in Video
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
Diffusion Transformers with Representation Autoencoders
von: Zheng, Boyang, et al.
Veröffentlicht: (2025)
von: Zheng, Boyang, et al.
Veröffentlicht: (2025)
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
Fast and Exact Enumeration of Deep Networks Partitions Regions
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
von: Dawid, Anna, et al.
Veröffentlicht: (2023)
von: Dawid, Anna, et al.
Veröffentlicht: (2023)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
Learning and Leveraging World Models in Visual Representation Learning
von: Garrido, Quentin, et al.
Veröffentlicht: (2024)
von: Garrido, Quentin, et al.
Veröffentlicht: (2024)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
von: Han, Junlin, et al.
Veröffentlicht: (2025)
von: Han, Junlin, et al.
Veröffentlicht: (2025)
Stochastic positional embeddings improve masked image modeling
von: Bar, Amir, et al.
Veröffentlicht: (2023)
von: Bar, Amir, et al.
Veröffentlicht: (2023)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
PaintBench: Deterministic Evaluation of Precise Visual Editing
von: Xu, Kai, et al.
Veröffentlicht: (2026)
von: Xu, Kai, et al.
Veröffentlicht: (2026)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
von: Zhai, Yuexiang, et al.
Veröffentlicht: (2024)
von: Zhai, Yuexiang, et al.
Veröffentlicht: (2024)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
von: Huang, Hai, et al.
Veröffentlicht: (2025)
von: Huang, Hai, et al.
Veröffentlicht: (2025)
Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA
von: Huang, Hai, et al.
Veröffentlicht: (2026)
von: Huang, Hai, et al.
Veröffentlicht: (2026)
Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
von: Dupoux, Emmanuel, et al.
Veröffentlicht: (2026)
von: Dupoux, Emmanuel, et al.
Veröffentlicht: (2026)
Variance Covariance Regularization Enforces Pairwise Independence in Self-Supervised Representations
von: Mialon, Grégoire, et al.
Veröffentlicht: (2022)
von: Mialon, Grégoire, et al.
Veröffentlicht: (2022)
A hierarchical loss and its problems when classifying non-hierarchically
von: Wu, Cinna, et al.
Veröffentlicht: (2017)
von: Wu, Cinna, et al.
Veröffentlicht: (2017)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
von: Hu, Yushi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Language-Free Visual Representation Learning
von: Fan, David, et al.
Veröffentlicht: (2025) -
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
von: Balestriero, Randall, et al.
Veröffentlicht: (2025) -
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
von: Tong, Shengbang, et al.
Veröffentlicht: (2024) -
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
von: Tong, Shengbang, et al.
Veröffentlicht: (2026) -
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2026)