Frame-Level Internal Tool Use for Temporal Grounding in Audio LMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Joesph, Keung, Phillip, Wang, Jiaqi, Ahia, Orevaoghene, Smith, Noah A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
von: Pahwa, Ramit, et al.
Veröffentlicht: (2026)
von: Pahwa, Ramit, et al.
Veröffentlicht: (2026)
SoundReactor: Frame-level Online Video-to-Audio Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025)
von: Primus, Paul, et al.
Veröffentlicht: (2025)
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
FLToP CTC: Frame-Level Token Pruning via Relative Threshold for Efficient and Memory-Saving Decoding on Diverse Platforms
von: Shree, Atul, et al.
Veröffentlicht: (2025)
von: Shree, Atul, et al.
Veröffentlicht: (2025)
A2SB: Audio-to-Audio Schrodinger Bridges
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
Multi-bit Audio Watermarking
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
The Rarity of Musical Audio Signals Within the Space of Possible Audio Generation
von: Collins, Nick
Veröffentlicht: (2024)
von: Collins, Nick
Veröffentlicht: (2024)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
von: Long, Phillip, et al.
Veröffentlicht: (2026)
von: Long, Phillip, et al.
Veröffentlicht: (2026)
SNAC: Multi-Scale Neural Audio Codec
von: Siuzdak, Hubert, et al.
Veröffentlicht: (2024)
von: Siuzdak, Hubert, et al.
Veröffentlicht: (2024)
Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
Audio Editing with Non-Rigid Text Prompts
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
Dissecting Temporal Understanding in Text-to-Audio Retrieval
von: Oncescu, Andreea-Maria, et al.
Veröffentlicht: (2024)
von: Oncescu, Andreea-Maria, et al.
Veröffentlicht: (2024)
Unsupervised Composable Representations for Audio
von: Bindi, Giovanni, et al.
Veröffentlicht: (2024)
von: Bindi, Giovanni, et al.
Veröffentlicht: (2024)
Diffusion Models for Audio Restoration
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
Instabilities in Convnets for Raw Audio
von: Haider, Daniel, et al.
Veröffentlicht: (2023)
von: Haider, Daniel, et al.
Veröffentlicht: (2023)
Automatic Contextual Audio Denoising
von: Luong, Diep, et al.
Veröffentlicht: (2026)
von: Luong, Diep, et al.
Veröffentlicht: (2026)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
StethoLM: Audio Language Model for Cardiopulmonary Analysis Across Clinical Tasks
von: Wang, Yishan, et al.
Veröffentlicht: (2026)
von: Wang, Yishan, et al.
Veröffentlicht: (2026)
A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
ADD 2022: the First Audio Deep Synthesis Detection Challenge
von: Yi, Jiangyan, et al.
Veröffentlicht: (2022)
von: Yi, Jiangyan, et al.
Veröffentlicht: (2022)
I Guess That's Why They Call it the Blues: Causal Analysis for Audio Classifiers
von: Kelly, David A., et al.
Veröffentlicht: (2026)
von: Kelly, David A., et al.
Veröffentlicht: (2026)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
Audio Decoding by Inverse Problem Solving
von: T., Pedro J. Villasana, et al.
Veröffentlicht: (2024)
von: T., Pedro J. Villasana, et al.
Veröffentlicht: (2024)
Does Audio Deepfake Detection Generalize?
von: Müller, Nicolas M., et al.
Veröffentlicht: (2022)
von: Müller, Nicolas M., et al.
Veröffentlicht: (2022)
Aligning Audio Captions with Human Preferences
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
Variable Bitrate Residual Vector Quantization for Audio Coding
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
von: Pahwa, Ramit, et al.
Veröffentlicht: (2026) -
SoundReactor: Frame-level Online Video-to-Audio Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2025) -
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025) -
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024) -
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)