Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Haoyun, Xiao, Xin, Zhong, Jiang, Tian, Yu, Xiaohua, Dong, Mao, Yu, Wu, Hao, Wei, Kaiwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart Speakers
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
von: Chen, Yu, et al.
Veröffentlicht: (2025)
von: Chen, Yu, et al.
Veröffentlicht: (2025)
Identifying Hearing Difficulty Moments in Conversational Audio
von: Collins, Jack, et al.
Veröffentlicht: (2025)
von: Collins, Jack, et al.
Veröffentlicht: (2025)
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
von: Ren, Yong, et al.
Veröffentlicht: (2025)
von: Ren, Yong, et al.
Veröffentlicht: (2025)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers
von: Lin, Liang, et al.
Veröffentlicht: (2025)
von: Lin, Liang, et al.
Veröffentlicht: (2025)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
von: Ren, Yong, et al.
Veröffentlicht: (2024)
von: Ren, Yong, et al.
Veröffentlicht: (2024)
I Can Hear You: Selective Robust Training for Deepfake Audio Detection
von: Zhang, Zirui, et al.
Veröffentlicht: (2024)
von: Zhang, Zirui, et al.
Veröffentlicht: (2024)
Video-to-Audio Generation with Fine-grained Temporal Semantics
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
AudioGenie-Reasoner: A Training-Free Multi-Agent Framework for Coarse-to-Fine Audio Deep Reasoning
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
von: Ledder, Wessel, et al.
Veröffentlicht: (2024)
von: Ledder, Wessel, et al.
Veröffentlicht: (2024)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
Naturalistic Music Decoding from EEG Data via Latent Diffusion Models
von: Postolache, Emilian, et al.
Veröffentlicht: (2024)
von: Postolache, Emilian, et al.
Veröffentlicht: (2024)
Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
von: Chen, Yu-Wen, et al.
Veröffentlicht: (2025)
von: Chen, Yu-Wen, et al.
Veröffentlicht: (2025)
WAKE: Watermarking Audio with Key Enrichment
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
Compositional Audio Representation Learning
von: Sridhar, Sripathi, et al.
Veröffentlicht: (2024)
von: Sridhar, Sripathi, et al.
Veröffentlicht: (2024)
STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
von: Yao, Dong, et al.
Veröffentlicht: (2023)
von: Yao, Dong, et al.
Veröffentlicht: (2023)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
Audio Fingerprinting with Holographic Reduced Representations
von: Fujita, Yusuke, et al.
Veröffentlicht: (2024)
von: Fujita, Yusuke, et al.
Veröffentlicht: (2024)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
von: Wang, Linqin, et al.
Veröffentlicht: (2024)
von: Wang, Linqin, et al.
Veröffentlicht: (2024)
Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
von: Martinez-Lucas, Luz, et al.
Veröffentlicht: (2026)
von: Martinez-Lucas, Luz, et al.
Veröffentlicht: (2026)
Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
HearSmoking: Smoking Detection in Driving Environment via Acoustic Sensing on Smartphones
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
Natural Language Supervision for General-Purpose Audio Representations
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
von: Zhao, Hang, et al.
Veröffentlicht: (2024)
von: Zhao, Hang, et al.
Veröffentlicht: (2024)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Frequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection
von: Liao, Yuan, et al.
Veröffentlicht: (2025)
von: Liao, Yuan, et al.
Veröffentlicht: (2025)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
von: Tian, Haokun, et al.
Veröffentlicht: (2025) -
HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart Speakers
von: Xie, Yadong, et al.
Veröffentlicht: (2025) -
Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
von: Xin, Yifei, et al.
Veröffentlicht: (2024) -
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025) -
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
von: Chen, Yu, et al.
Veröffentlicht: (2025)