Thinking While Listening: Simple Test Time Scaling For Audio Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Verma, Prateek, Pilanci, Mert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Large Language Models By Layerwise Attention Shortcuts
by: Verma, Prateek, et al.
Published: (2024)
by: Verma, Prateek, et al.
Published: (2024)
Towards Signal Processing In Large Language Models
by: Verma, Prateek, et al.
Published: (2024)
by: Verma, Prateek, et al.
Published: (2024)
Large Language Models Implicitly Learn to See and Hear Just By Reading
by: Verma, Prateek, et al.
Published: (2025)
by: Verma, Prateek, et al.
Published: (2025)
Content Adaptive Front End For Audio Classification
by: Verma, Prateek, et al.
Published: (2023)
by: Verma, Prateek, et al.
Published: (2023)
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
by: Verma, Prateek
Published: (2023)
by: Verma, Prateek
Published: (2023)
Audio Transformers
by: Verma, Prateek, et al.
Published: (2021)
by: Verma, Prateek, et al.
Published: (2021)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
by: Verma, Prateek
Published: (2024)
by: Verma, Prateek
Published: (2024)
A Language Model With Million Context Length For Raw Audio
by: Verma, Prateek
Published: (2022)
by: Verma, Prateek
Published: (2022)
AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers
by: Biju, Emil, et al.
Published: (2024)
by: Biju, Emil, et al.
Published: (2024)
Neural Style Transfer for Audio Spectograms
by: Verma, Prateek, et al.
Published: (2018)
by: Verma, Prateek, et al.
Published: (2018)
Wavelet GPT: Wavelet Inspired Large Language Models
by: Verma, Prateek
Published: (2024)
by: Verma, Prateek
Published: (2024)
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
by: Du, Zhihao, et al.
Published: (2023)
by: Du, Zhihao, et al.
Published: (2023)
$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
by: Maharana, Sarthak Kumar, et al.
Published: (2025)
by: Maharana, Sarthak Kumar, et al.
Published: (2025)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
by: Zhang, Wenda, et al.
Published: (2026)
by: Zhang, Wenda, et al.
Published: (2026)
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
by: Jia, Hong, et al.
Published: (2026)
by: Jia, Hong, et al.
Published: (2026)
Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
by: Işık, Atakan, et al.
Published: (2025)
by: Işık, Atakan, et al.
Published: (2025)
Guiding Audio Editing with Audio Language Model
by: Lan, Zitong, et al.
Published: (2025)
by: Lan, Zitong, et al.
Published: (2025)
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
by: Dang, Ting, et al.
Published: (2025)
by: Dang, Ting, et al.
Published: (2025)
AudioGenX: Explainability on Text-to-Audio Generative Models
by: Kang, Hyunju, et al.
Published: (2025)
by: Kang, Hyunju, et al.
Published: (2025)
Simple and Controllable Music Generation
by: Copet, Jade, et al.
Published: (2023)
by: Copet, Jade, et al.
Published: (2023)
HEAR: Holistic Evaluation of Audio Representations
by: Turian, Joseph, et al.
Published: (2022)
by: Turian, Joseph, et al.
Published: (2022)
Listenable Maps for Audio Classifiers
by: Paissan, Francesco, et al.
Published: (2024)
by: Paissan, Francesco, et al.
Published: (2024)
Learning Source Disentanglement in Neural Audio Codec
by: Bie, Xiaoyu, et al.
Published: (2024)
by: Bie, Xiaoyu, et al.
Published: (2024)
TACNET: Temporal Audio Source Counting Network
by: Ahmadnejad, Amirreza, et al.
Published: (2023)
by: Ahmadnejad, Amirreza, et al.
Published: (2023)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
by: Mancini, Eleonora, et al.
Published: (2024)
by: Mancini, Eleonora, et al.
Published: (2024)
Training-Free Multi-Step Audio Source Separation
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
On Temporal Guidance and Iterative Refinement in Audio Source Separation
by: Morocutti, Tobias, et al.
Published: (2025)
by: Morocutti, Tobias, et al.
Published: (2025)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)
by: Mancini, Eleonora, et al.
Published: (2025)
Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Exploring and Applying Audio-Based Sentiment Analysis in Music
by: Jhanji, Etash
Published: (2024)
by: Jhanji, Etash
Published: (2024)
Prompt-guided Precise Audio Editing with Diffusion Models
by: Xu, Manjie, et al.
Published: (2024)
by: Xu, Manjie, et al.
Published: (2024)
A Survey of Deep Learning Audio Generation Methods
by: Božić, Matej, et al.
Published: (2024)
by: Božić, Matej, et al.
Published: (2024)
Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
by: Torres, Bernardo, et al.
Published: (2025)
by: Torres, Bernardo, et al.
Published: (2025)
Text-Queried Audio Source Separation via Hierarchical Modeling
by: Yin, Xinlei, et al.
Published: (2025)
by: Yin, Xinlei, et al.
Published: (2025)
Sparse Autoencoders Make Audio Foundation Models more Explainable
by: Mariotte, Théo, et al.
Published: (2025)
by: Mariotte, Théo, et al.
Published: (2025)
Evaluating Fake Music Detection Performance Under Audio Augmentations
by: Sroka, Tomasz, et al.
Published: (2025)
by: Sroka, Tomasz, et al.
Published: (2025)
Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)
by: Long, Phillip, et al.
Published: (2026)
PoDAR: Power-Disentangled Audio Representation for Generative Modeling
by: Luebs, Alejandro, et al.
Published: (2026)
by: Luebs, Alejandro, et al.
Published: (2026)
Similar Items
-
Adaptive Large Language Models By Layerwise Attention Shortcuts
by: Verma, Prateek, et al.
Published: (2024) -
Towards Signal Processing In Large Language Models
by: Verma, Prateek, et al.
Published: (2024) -
Large Language Models Implicitly Learn to See and Hear Just By Reading
by: Verma, Prateek, et al.
Published: (2025) -
Content Adaptive Front End For Audio Classification
by: Verma, Prateek, et al.
Published: (2023) -
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
by: Verma, Prateek
Published: (2023)