LLMs can see and hear without any training
Fuente:
arXiv
Saved in:
| Main Authors: | Ashutosh, Kumar, Gandelsman, Yossi, Chen, Xinlei, Misra, Ishan, Girdhar, Rohit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023)
by: Menon, Sachit, et al.
Published: (2023)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Learning Video Representations without Natural Videos
by: Yu, Xueyang, et al.
Published: (2024)
by: Yu, Xueyang, et al.
Published: (2024)
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
by: Bu, Edmund, et al.
Published: (2025)
by: Bu, Edmund, et al.
Published: (2025)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
Transformers without Normalization
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
An Empirical Study of Autoregressive Pre-training from Videos
by: Rajasegaran, Jathushan, et al.
Published: (2025)
by: Rajasegaran, Jathushan, et al.
Published: (2025)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
Jailbreaking Vision-Language Models Through the Visual Modality
by: Azulay, Aharon, et al.
Published: (2026)
by: Azulay, Aharon, et al.
Published: (2026)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
by: Singh, Mannat, et al.
Published: (2023)
by: Singh, Mannat, et al.
Published: (2023)
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
Human detectors are surprisingly powerful reward models
by: Ashutosh, Kumar, et al.
Published: (2026)
by: Ashutosh, Kumar, et al.
Published: (2026)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Synthesizing Moving People with 3D Control
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
by: Zang, Yuan, et al.
Published: (2024)
by: Zang, Yuan, et al.
Published: (2024)
Using Shapley interactions to understand how models use structure
by: Singhvi, Divyansh, et al.
Published: (2024)
by: Singhvi, Divyansh, et al.
Published: (2024)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
Physical Property Understanding from Language-Embedded Feature Fields
by: Zhai, Albert J., et al.
Published: (2024)
by: Zhai, Albert J., et al.
Published: (2024)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
by: Cai, Zhenyang, et al.
Published: (2024)
by: Cai, Zhenyang, et al.
Published: (2024)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
by: Lou, Siyu, et al.
Published: (2024)
by: Lou, Siyu, et al.
Published: (2024)
Is Pre-training Truly Better Than Meta-Learning?
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Efficient Pre-training for Localized Instruction Generation of Videos
by: Batra, Anil, et al.
Published: (2023)
by: Batra, Anil, et al.
Published: (2023)
iSign: A Benchmark for Indian Sign Language Processing
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
by: Chen, Lu, et al.
Published: (2024)
by: Chen, Lu, et al.
Published: (2024)
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
by: Li, Hong, et al.
Published: (2024)
by: Li, Hong, et al.
Published: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
by: Chen, Junying, et al.
Published: (2024)
by: Chen, Junying, et al.
Published: (2024)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
by: Ekin, Yigit, et al.
Published: (2026)
by: Ekin, Yigit, et al.
Published: (2026)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
by: Deng, Naihao, et al.
Published: (2024)
by: Deng, Naihao, et al.
Published: (2024)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
by: Javadi, Farnoosh, et al.
Published: (2023)
by: Javadi, Farnoosh, et al.
Published: (2023)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
by: Wan, David, et al.
Published: (2024)
by: Wan, David, et al.
Published: (2024)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
by: Fang, Xianya, et al.
Published: (2026)
by: Fang, Xianya, et al.
Published: (2026)
Similar Items
-
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023) -
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025) -
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024) -
Learning Video Representations without Natural Videos
by: Yu, Xueyang, et al.
Published: (2024) -
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
by: Bu, Edmund, et al.
Published: (2025)