Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Tong, Shengbang, Liu, Zhuang, Zhai, Yuexiang, Ma, Yi, LeCun, Yann, Xie, Saining |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
por: Tong, Shengbang, et al.
Publicado: (2024)
por: Tong, Shengbang, et al.
Publicado: (2024)
Scaling Language-Free Visual Representation Learning
por: Fan, David, et al.
Publicado: (2025)
por: Fan, David, et al.
Publicado: (2025)
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
por: Zhai, Yuexiang, et al.
Publicado: (2024)
por: Zhai, Yuexiang, et al.
Publicado: (2024)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
por: Tong, Shengbang, et al.
Publicado: (2026)
por: Tong, Shengbang, et al.
Publicado: (2026)
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
por: Tong, Shengbang, et al.
Publicado: (2024)
por: Tong, Shengbang, et al.
Publicado: (2024)
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
por: Zhou, Yikang, et al.
Publicado: (2025)
por: Zhou, Yikang, et al.
Publicado: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
por: Chu, Tianzhe, et al.
Publicado: (2025)
por: Chu, Tianzhe, et al.
Publicado: (2025)
Diffusion Transformers with Representation Autoencoders
por: Zheng, Boyang, et al.
Publicado: (2025)
por: Zheng, Boyang, et al.
Publicado: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
por: Balestriero, Randall, et al.
Publicado: (2025)
por: Balestriero, Randall, et al.
Publicado: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
por: Balestriero, Randall, et al.
Publicado: (2024)
por: Balestriero, Randall, et al.
Publicado: (2024)
A hierarchical loss and its problems when classifying non-hierarchically
por: Wu, Cinna, et al.
Publicado: (2017)
por: Wu, Cinna, et al.
Publicado: (2017)
Video Representation Learning with Joint-Embedding Predictive Architectures
por: Drozdov, Katrina, et al.
Publicado: (2024)
por: Drozdov, Katrina, et al.
Publicado: (2024)
Cambrian-S: Towards Spatial Supersensing in Video
por: Yang, Shusheng, et al.
Publicado: (2025)
por: Yang, Shusheng, et al.
Publicado: (2025)
The Entropy Enigma: Success and Failure of Entropy Minimization
por: Press, Ori, et al.
Publicado: (2024)
por: Press, Ori, et al.
Publicado: (2024)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
por: Tong, Shengbang, et al.
Publicado: (2026)
por: Tong, Shengbang, et al.
Publicado: (2026)
URLOST: Unsupervised Representation Learning without Stationarity or Topology
por: Yun, Zeyu, et al.
Publicado: (2023)
por: Yun, Zeyu, et al.
Publicado: (2023)
Transformers without Normalization
por: Zhu, Jiachen, et al.
Publicado: (2025)
por: Zhu, Jiachen, et al.
Publicado: (2025)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
por: Goswami, Raktim Gautam, et al.
Publicado: (2024)
por: Goswami, Raktim Gautam, et al.
Publicado: (2024)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
por: Hansen, Nicklas, et al.
Publicado: (2024)
por: Hansen, Nicklas, et al.
Publicado: (2024)
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density
por: Balestriero, Randall, et al.
Publicado: (2025)
por: Balestriero, Randall, et al.
Publicado: (2025)
Learning and Leveraging World Models in Visual Representation Learning
por: Garrido, Quentin, et al.
Publicado: (2024)
por: Garrido, Quentin, et al.
Publicado: (2024)
Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
por: Denton, Remi, et al.
Publicado: (2014)
por: Denton, Remi, et al.
Publicado: (2014)
Blockwise Self-Supervised Learning at Scale
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2023)
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2023)
Asymmetric Idiosyncrasies in Multimodal Models
por: Tao, Muzi, et al.
Publicado: (2026)
por: Tao, Muzi, et al.
Publicado: (2026)
PooDLe: Pooled and dense self-supervised learning from naturalistic videos
por: Wang, Alex N., et al.
Publicado: (2024)
por: Wang, Alex N., et al.
Publicado: (2024)
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
por: Rudman, William, et al.
Publicado: (2025)
por: Rudman, William, et al.
Publicado: (2025)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
por: Chandhok, Shivam, et al.
Publicado: (2024)
por: Chandhok, Shivam, et al.
Publicado: (2024)
Learning Latent Action World Models In The Wild
por: Garrido, Quentin, et al.
Publicado: (2026)
por: Garrido, Quentin, et al.
Publicado: (2026)
Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
por: Kuang, Yilun, et al.
Publicado: (2026)
por: Kuang, Yilun, et al.
Publicado: (2026)
Navigation World Models
por: Bar, Amir, et al.
Publicado: (2024)
por: Bar, Amir, et al.
Publicado: (2024)
Variance-Covariance Regularization Improves Representation Learning
por: Zhu, Jiachen, et al.
Publicado: (2023)
por: Zhu, Jiachen, et al.
Publicado: (2023)
Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimation
por: Zeevi, Tal, et al.
Publicado: (2024)
por: Zeevi, Tal, et al.
Publicado: (2024)
Revisiting Feature Prediction for Learning Visual Representations from Video
por: Bardes, Adrien, et al.
Publicado: (2024)
por: Bardes, Adrien, et al.
Publicado: (2024)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
por: Yeh, Chun-Hsiao, et al.
Publicado: (2025)
por: Yeh, Chun-Hsiao, et al.
Publicado: (2025)
Improving Pre-trained Self-Supervised Embeddings Through Effective Entropy Maximization
por: Chakraborty, Deep, et al.
Publicado: (2024)
por: Chakraborty, Deep, et al.
Publicado: (2024)
Eyes Tell the Truth: GazeVal Highlights Shortcomings of Generative AI in Medical Imaging
por: Wong, David, et al.
Publicado: (2025)
por: Wong, David, et al.
Publicado: (2025)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
por: Zhou, Guanyu, et al.
Publicado: (2026)
por: Zhou, Guanyu, et al.
Publicado: (2026)
Eyes Will Shut: A Vision-Based Next GPS Location Prediction Model by Reinforcement Learning from Visual Map Feed Back
por: Zhang, Ruixing, et al.
Publicado: (2025)
por: Zhang, Ruixing, et al.
Publicado: (2025)
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
por: Yu, Yaodong, et al.
Publicado: (2023)
por: Yu, Yaodong, et al.
Publicado: (2023)
Flow Map Distillation Without Data
por: Tong, Shangyuan, et al.
Publicado: (2025)
por: Tong, Shangyuan, et al.
Publicado: (2025)
Ejemplares similares
-
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
por: Tong, Shengbang, et al.
Publicado: (2024) -
Scaling Language-Free Visual Representation Learning
por: Fan, David, et al.
Publicado: (2025) -
Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
por: Zhai, Yuexiang, et al.
Publicado: (2024) -
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
por: Tong, Shengbang, et al.
Publicado: (2026) -
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
por: Tong, Shengbang, et al.
Publicado: (2024)