Large Language Models Implicitly Learn to See and Hear Just By Reading
Fuente:
arXiv
Saved in:
| Main Authors: | Verma, Prateek, Pilanci, Mert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Signal Processing In Large Language Models
by: Verma, Prateek, et al.
Published: (2024)
by: Verma, Prateek, et al.
Published: (2024)
Adaptive Large Language Models By Layerwise Attention Shortcuts
by: Verma, Prateek, et al.
Published: (2024)
by: Verma, Prateek, et al.
Published: (2024)
Thinking While Listening: Simple Test Time Scaling For Audio Classification
by: Verma, Prateek, et al.
Published: (2025)
by: Verma, Prateek, et al.
Published: (2025)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
by: Verma, Prateek
Published: (2024)
by: Verma, Prateek
Published: (2024)
Wavelet GPT: Wavelet Inspired Large Language Models
by: Verma, Prateek
Published: (2024)
by: Verma, Prateek
Published: (2024)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
Tell What You Hear From What You See -- Video to Audio Generation Through Text
by: Liu, Xiulong, et al.
Published: (2024)
by: Liu, Xiulong, et al.
Published: (2024)
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
by: Chen, Mingfei, et al.
Published: (2025)
by: Chen, Mingfei, et al.
Published: (2025)
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
by: Zhang, Shaolei, et al.
Published: (2025)
by: Zhang, Shaolei, et al.
Published: (2025)
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
by: Jia, Yanhao, et al.
Published: (2025)
by: Jia, Yanhao, et al.
Published: (2025)
Hearing Anywhere in Any Environment
by: Liu, Xiulong, et al.
Published: (2025)
by: Liu, Xiulong, et al.
Published: (2025)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
by: Tang, Yiwen, et al.
Published: (2024)
by: Tang, Yiwen, et al.
Published: (2024)
Towards Multi-Modal Mastery: A 4.5B Parameter Truly Multi-Modal Small Language Model
by: Koska, Ben, et al.
Published: (2024)
by: Koska, Ben, et al.
Published: (2024)
Learning Separable Hidden Unit Contributions for Speaker-Adaptive Lip-Reading
by: Luo, Songtao, et al.
Published: (2023)
by: Luo, Songtao, et al.
Published: (2023)
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
by: Tang, Mingni, et al.
Published: (2025)
by: Tang, Mingni, et al.
Published: (2025)
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
by: Park, Sooyoung, et al.
Published: (2025)
by: Park, Sooyoung, et al.
Published: (2025)
Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learning
by: Cheng, Xin, et al.
Published: (2025)
by: Cheng, Xin, et al.
Published: (2025)
The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
by: He, Luxi, et al.
Published: (2025)
by: He, Luxi, et al.
Published: (2025)
Ming-Omni: A Unified Multimodal Model for Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
by: Wang, Jinting, et al.
Published: (2025)
by: Wang, Jinting, et al.
Published: (2025)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
by: Rong, Yan, et al.
Published: (2024)
by: Rong, Yan, et al.
Published: (2024)
What Do Language Models Hear? Probing for Auditory Representations in Language Models
by: Ngo, Jerry, et al.
Published: (2024)
by: Ngo, Jerry, et al.
Published: (2024)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
by: Tan, Weiting, et al.
Published: (2025)
by: Tan, Weiting, et al.
Published: (2025)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
by: Liang, Yingshan, et al.
Published: (2025)
by: Liang, Yingshan, et al.
Published: (2025)
Automated Extraction of Spatio-Semantic Graphs for Identifying Cognitive Impairment
by: Ng, Si-Ioi, et al.
Published: (2025)
by: Ng, Si-Ioi, et al.
Published: (2025)
MMFformer: Multimodal Fusion Transformer Network for Depression Detection
by: Haque, Md Rezwanul, et al.
Published: (2025)
by: Haque, Md Rezwanul, et al.
Published: (2025)
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
by: Wang, Zhongxi, et al.
Published: (2026)
by: Wang, Zhongxi, et al.
Published: (2026)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
Imagine to Hear: Auditory Knowledge Generation can be an Effective Assistant for Language Models
by: Yoo, Suho, et al.
Published: (2025)
by: Yoo, Suho, et al.
Published: (2025)
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
by: Liu, Zhi-Song, et al.
Published: (2024)
by: Liu, Zhi-Song, et al.
Published: (2024)
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
by: Kang, Minjae, et al.
Published: (2025)
by: Kang, Minjae, et al.
Published: (2025)
SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
by: Wang, Kaidi, et al.
Published: (2025)
by: Wang, Kaidi, et al.
Published: (2025)
Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition
by: Deng, Shijian, et al.
Published: (2024)
by: Deng, Shijian, et al.
Published: (2024)
Spoken Language Intelligence of Large Language Models for Language Learning
by: Peng, Linkai, et al.
Published: (2023)
by: Peng, Linkai, et al.
Published: (2023)
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
by: Namgyal, Tashi, et al.
Published: (2024)
by: Namgyal, Tashi, et al.
Published: (2024)
Similar Items
-
Towards Signal Processing In Large Language Models
by: Verma, Prateek, et al.
Published: (2024) -
Adaptive Large Language Models By Layerwise Attention Shortcuts
by: Verma, Prateek, et al.
Published: (2024) -
Thinking While Listening: Simple Test Time Scaling For Audio Classification
by: Verma, Prateek, et al.
Published: (2025) -
Whisper-GPT: A Hybrid Representation Audio Large Language Model
by: Verma, Prateek
Published: (2024) -
Wavelet GPT: Wavelet Inspired Large Language Models
by: Verma, Prateek
Published: (2024)