Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
Fuente:
arXiv
Saved in:
| Main Authors: | Loakman, Tyler, James, Joseph, Lin, Chenghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
by: Loakman, Tyler, et al.
Published: (2024)
by: Loakman, Tyler, et al.
Published: (2024)
ReproHum #0087-01: Human Evaluation Reproduction Report for Generating Fact Checking Explanations
by: Loakman, Tyler, et al.
Published: (2024)
by: Loakman, Tyler, et al.
Published: (2024)
Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
by: Loakman, Tyler, et al.
Published: (2025)
by: Loakman, Tyler, et al.
Published: (2025)
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes
by: Loakman, Tyler, et al.
Published: (2025)
by: Loakman, Tyler, et al.
Published: (2025)
Train & Constrain: Phonologically Informed Tongue-Twister Generation from Topics and Paraphrases
by: Loakman, Tyler, et al.
Published: (2024)
by: Loakman, Tyler, et al.
Published: (2024)
Is neural semantic parsing good at ellipsis resolution, or isn't it?
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Hearing what isn't said
by: Yasuko Maeda
Published: (2024)
by: Yasuko Maeda
Published: (2024)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
by: Wang, Shun, et al.
Published: (2024)
by: Wang, Shun, et al.
Published: (2024)
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
by: Wang, Shun, et al.
Published: (2025)
by: Wang, Shun, et al.
Published: (2025)
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
CADGE: Context-Aware Dialogue Generation Enhanced with Graph-Structured Knowledge Aggregation
by: Zhang, Hongbo, et al.
Published: (2023)
by: Zhang, Hongbo, et al.
Published: (2023)
Moreover. Shakespeare it isn't
Published: (1998)
Published: (1998)
Why isn’t Mexico on China’s Growth Path?
by: James Gerber
Published: (2012)
by: James Gerber
Published: (2012)
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
Rampant caries: What it is and what it isn't
by: Mahen Ganhewa, et al.
Published: (2024)
by: Mahen Ganhewa, et al.
Published: (2024)
Bill gates isn't in the clear yet
Poverty traps are rare, but trappedness isn't
by: Mengesha, Isaak, et al.
Published: (2026)
by: Mengesha, Isaak, et al.
Published: (2026)
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
by: Wu, Siwei, et al.
Published: (2025)
by: Wu, Siwei, et al.
Published: (2025)
Finding Challenging Metaphors that Confuse Pretrained Language Models
by: Li, Yucheng, et al.
Published: (2024)
by: Li, Yucheng, et al.
Published: (2024)
United States. Anthrax isn't contagious: nxiety is
Published: (2001)
Published: (2001)
United states. Alabama isn't so different
Published: (1997)
Published: (1997)
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights
by: James, Joseph, et al.
Published: (2024)
by: James, Joseph, et al.
Published: (2024)
Natural Language Generation
by: van Miltenburg, Emiel, et al.
Published: (2025)
by: van Miltenburg, Emiel, et al.
Published: (2025)
Leveraging Large Language Models for Zero-shot Lay Summarisation in Biomedicine and Beyond
by: Goldsack, Tomas, et al.
Published: (2025)
by: Goldsack, Tomas, et al.
Published: (2025)
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
by: Sun, Yizheng, et al.
Published: (2025)
by: Sun, Yizheng, et al.
Published: (2025)
An Open Source Data Contamination Report for Large Language Models
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
by: Wang, Jin, et al.
Published: (2026)
by: Wang, Jin, et al.
Published: (2026)
Life at low Reynolds number isn't such a drag
by: Datta, Sujit S.
Published: (2024)
by: Datta, Sujit S.
Published: (2024)
Hotter isn't faster for a melting RNA hairpin
by: Li, Huaping, et al.
Published: (2024)
by: Li, Huaping, et al.
Published: (2024)
Language Model as an Annotator: Unsupervised Context-aware Quality Phrase Generation
by: Zhang, Zhihao, et al.
Published: (2023)
by: Zhang, Zhihao, et al.
Published: (2023)
Large Language Models Implicitly Learn to See and Hear Just By Reading
by: Verma, Prateek, et al.
Published: (2025)
by: Verma, Prateek, et al.
Published: (2025)
LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models
by: Li, Haokun, et al.
Published: (2025)
by: Li, Haokun, et al.
Published: (2025)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026)
by: James, Joseph, et al.
Published: (2026)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
by: Yeh, Sung-Lin, et al.
Published: (2026)
by: Yeh, Sung-Lin, et al.
Published: (2026)
Evaluating Large Language Models for Generalization and Robustness via Data Compression
by: Li, Yucheng, et al.
Published: (2024)
by: Li, Yucheng, et al.
Published: (2024)
Evidence Generation for Drugs and Biological Products isn't Magic or Myth
by: John Concato, et al.
Published: (2025)
by: John Concato, et al.
Published: (2025)
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Similar Items
-
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
by: Loakman, Tyler, et al.
Published: (2024) -
ReproHum #0087-01: Human Evaluation Reproduction Report for Generating Fact Checking Explanations
by: Loakman, Tyler, et al.
Published: (2024) -
Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
by: Loakman, Tyler, et al.
Published: (2025) -
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes
by: Loakman, Tyler, et al.
Published: (2025) -
Train & Constrain: Phonologically Informed Tongue-Twister Generation from Topics and Paraphrases
by: Loakman, Tyler, et al.
Published: (2024)