Scaling Language-Free Visual Representation Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Fan, David, Tong, Shengbang, Zhu, Jiachen, Sinha, Koustuv, Liu, Zhuang, Chen, Xinlei, Rabbat, Michael, Ballas, Nicolas, LeCun, Yann, Bar, Amir, Xie, Saining |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
por: Tong, Shengbang, et al.
Publicado: (2024)
por: Tong, Shengbang, et al.
Publicado: (2024)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
por: Balestriero, Randall, et al.
Publicado: (2025)
por: Balestriero, Randall, et al.
Publicado: (2025)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
por: Tong, Shengbang, et al.
Publicado: (2024)
por: Tong, Shengbang, et al.
Publicado: (2024)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
por: Mur-Labadia, Lorenzo, et al.
Publicado: (2026)
por: Mur-Labadia, Lorenzo, et al.
Publicado: (2026)
Revisiting Feature Prediction for Learning Visual Representations from Video
por: Bardes, Adrien, et al.
Publicado: (2024)
por: Bardes, Adrien, et al.
Publicado: (2024)
Parallel Stochastic Gradient-Based Planning for World Models
por: Psenka, Michael, et al.
Publicado: (2026)
por: Psenka, Michael, et al.
Publicado: (2026)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
por: Tong, Shengbang, et al.
Publicado: (2026)
por: Tong, Shengbang, et al.
Publicado: (2026)
Learning Latent Action World Models In The Wild
por: Garrido, Quentin, et al.
Publicado: (2026)
por: Garrido, Quentin, et al.
Publicado: (2026)
Learning and Leveraging World Models in Visual Representation Learning
por: Garrido, Quentin, et al.
Publicado: (2024)
por: Garrido, Quentin, et al.
Publicado: (2024)
Transformers without Normalization
por: Zhu, Jiachen, et al.
Publicado: (2025)
por: Zhu, Jiachen, et al.
Publicado: (2025)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
por: Tong, Shengbang, et al.
Publicado: (2026)
por: Tong, Shengbang, et al.
Publicado: (2026)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
por: Terver, Basile, et al.
Publicado: (2026)
por: Terver, Basile, et al.
Publicado: (2026)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
por: Garrido, Quentin, et al.
Publicado: (2025)
por: Garrido, Quentin, et al.
Publicado: (2025)
Stochastic positional embeddings improve masked image modeling
por: Bar, Amir, et al.
Publicado: (2023)
por: Bar, Amir, et al.
Publicado: (2023)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
por: Sun, Shangwen, et al.
Publicado: (2026)
por: Sun, Shangwen, et al.
Publicado: (2026)
Variance-Covariance Regularization Improves Representation Learning
por: Zhu, Jiachen, et al.
Publicado: (2023)
por: Zhu, Jiachen, et al.
Publicado: (2023)
Diffusion Transformers with Representation Autoencoders
por: Zheng, Boyang, et al.
Publicado: (2025)
por: Zheng, Boyang, et al.
Publicado: (2025)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
por: Zhai, Yuexiang, et al.
Publicado: (2024)
por: Zhai, Yuexiang, et al.
Publicado: (2024)
Navigation World Models
por: Bar, Amir, et al.
Publicado: (2024)
por: Bar, Amir, et al.
Publicado: (2024)
Variance Covariance Regularization Enforces Pairwise Independence in Self-Supervised Representations
por: Mialon, Grégoire, et al.
Publicado: (2022)
por: Mialon, Grégoire, et al.
Publicado: (2022)
Fast and Exact Enumeration of Deep Networks Partitions Regions
por: Balestriero, Randall, et al.
Publicado: (2024)
por: Balestriero, Randall, et al.
Publicado: (2024)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
por: Dawid, Anna, et al.
Publicado: (2023)
por: Dawid, Anna, et al.
Publicado: (2023)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
por: Balestriero, Randall, et al.
Publicado: (2025)
por: Balestriero, Randall, et al.
Publicado: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
por: Balestriero, Randall, et al.
Publicado: (2024)
por: Balestriero, Randall, et al.
Publicado: (2024)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
por: Han, Junlin, et al.
Publicado: (2025)
por: Han, Junlin, et al.
Publicado: (2025)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
por: Goswami, Raktim Gautam, et al.
Publicado: (2025)
por: Goswami, Raktim Gautam, et al.
Publicado: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
por: Drozdov, Katrina, et al.
Publicado: (2024)
por: Drozdov, Katrina, et al.
Publicado: (2024)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
por: Huang, Hai, et al.
Publicado: (2025)
por: Huang, Hai, et al.
Publicado: (2025)
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
por: Rudman, William, et al.
Publicado: (2025)
por: Rudman, William, et al.
Publicado: (2025)
Whole-Body Conditioned Egocentric Video Prediction
por: Bai, Yutong, et al.
Publicado: (2025)
por: Bai, Yutong, et al.
Publicado: (2025)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
por: Yun, Zeyu, et al.
Publicado: (2023)
por: Yun, Zeyu, et al.
Publicado: (2023)
Hierarchical Planning with Latent World Models
por: Zhang, Wancong, et al.
Publicado: (2026)
por: Zhang, Wancong, et al.
Publicado: (2026)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
por: Skean, Oscar, et al.
Publicado: (2024)
por: Skean, Oscar, et al.
Publicado: (2024)
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
por: Zhou, Gaoyue, et al.
Publicado: (2024)
por: Zhou, Gaoyue, et al.
Publicado: (2024)
Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA
por: Huang, Hai, et al.
Publicado: (2026)
por: Huang, Hai, et al.
Publicado: (2026)
Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
por: Dupoux, Emmanuel, et al.
Publicado: (2026)
por: Dupoux, Emmanuel, et al.
Publicado: (2026)
A hierarchical loss and its problems when classifying non-hierarchically
por: Wu, Cinna, et al.
Publicado: (2017)
por: Wu, Cinna, et al.
Publicado: (2017)
Blockwise Self-Supervised Learning at Scale
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2023)
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2023)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
por: Lin, Han, et al.
Publicado: (2024)
por: Lin, Han, et al.
Publicado: (2024)
EgoPet: Egomotion and Interaction Data from an Animal's Perspective
por: Bar, Amir, et al.
Publicado: (2024)
por: Bar, Amir, et al.
Publicado: (2024)
Ejemplares similares
-
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
por: Tong, Shengbang, et al.
Publicado: (2024) -
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
por: Balestriero, Randall, et al.
Publicado: (2025) -
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
por: Tong, Shengbang, et al.
Publicado: (2024) -
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
por: Mur-Labadia, Lorenzo, et al.
Publicado: (2026) -
Revisiting Feature Prediction for Learning Visual Representations from Video
por: Bardes, Adrien, et al.
Publicado: (2024)