ADIFF: Explaining audio difference using natural language
Fuente:
Zenodo
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Recurso digital |
| Publié: |
Zenodo
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866902141049765888 |
|---|---|
| author | Deshmukh, Soham Han, Shuo Singh, Rita Raj, Bhiksha |
| author_facet | Deshmukh, Soham Han, Shuo Singh, Rita Raj, Bhiksha |
| contents | <div> <div>ADIFF is an audio prefix tuning-based language model with a cross-projection module and undergoes a three-step training process. ADIFF takes two audios and text prompt as input and produces different tiers of difference explanations as output. This involves identifying and describing audio events, acoustic scenes, signal characteristics, and their emotional impact on listeners.</div> <div><br>The code repository is: <a href="https://github.com/soham97/ADIFF">soham97/ADIFF</a></div> </div> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_14706090 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | ADIFF: Explaining audio difference using natural language Deshmukh, Soham Han, Shuo Singh, Rita Raj, Bhiksha audio-language multimodal learning audio question answering sound events acoustic scenes adiff zero-shot <div> <div>ADIFF is an audio prefix tuning-based language model with a cross-projection module and undergoes a three-step training process. ADIFF takes two audios and text prompt as input and produces different tiers of difference explanations as output. This involves identifying and describing audio events, acoustic scenes, signal characteristics, and their emotional impact on listeners.</div> <div><br>The code repository is: <a href="https://github.com/soham97/ADIFF">soham97/ADIFF</a></div> </div> |
| title | ADIFF: Explaining audio difference using natural language |
| topic | audio-language multimodal learning audio question answering sound events acoustic scenes adiff zero-shot |
| url | https://doi.org/10.5281/zenodo.14706090 |