Audio Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Verma, Prateek, Berger, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
von: Verma, Prateek
Veröffentlicht: (2023)
von: Verma, Prateek
Veröffentlicht: (2023)
Content Adaptive Front End For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
Neural Style Transfer for Audio Spectograms
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
Generative AI for Music and Audio
von: Dong, Hao-Wen
Veröffentlicht: (2024)
von: Dong, Hao-Wen
Veröffentlicht: (2024)
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Fast Text-to-Audio Generation with Adversarial Post-Training
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
von: Du, Zhihao, et al.
Veröffentlicht: (2023)
von: Du, Zhihao, et al.
Veröffentlicht: (2023)
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2025)
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2025)
Understanding Pedestrian Movement Using Urban Sensing Technologies: The Promise of Audio-based Sensors
von: Han, Chaeyeon, et al.
Veröffentlicht: (2024)
von: Han, Chaeyeon, et al.
Veröffentlicht: (2024)
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
Thinking While Listening: Simple Test Time Scaling For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
von: Boudaghi, Ali, et al.
Veröffentlicht: (2025)
von: Boudaghi, Ali, et al.
Veröffentlicht: (2025)
FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2024)
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2024)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
von: Li, Maomao, et al.
Veröffentlicht: (2026)
von: Li, Maomao, et al.
Veröffentlicht: (2026)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
von: Verma, Prateek
Veröffentlicht: (2024)
von: Verma, Prateek
Veröffentlicht: (2024)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
Embedding Alignment in Code Generation for Audio
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
Kimi-Audio Technical Report
von: KimiTeam, et al.
Veröffentlicht: (2025)
von: KimiTeam, et al.
Veröffentlicht: (2025)
Integrating IP Broadcasting with Audio Tags: Workflow and Challenges
von: Burchett-Vass, Rhys, et al.
Veröffentlicht: (2024)
von: Burchett-Vass, Rhys, et al.
Veröffentlicht: (2024)
Unveiling Visual Biases in Audio-Visual Localization Benchmarks
von: Chen, Liangyu, et al.
Veröffentlicht: (2024)
von: Chen, Liangyu, et al.
Veröffentlicht: (2024)
Automatic Music Transcription using Convolutional Neural Networks and Constant-Q transform
von: Telila, Yohannis, et al.
Veröffentlicht: (2025)
von: Telila, Yohannis, et al.
Veröffentlicht: (2025)
Sequence-to-Sequence Multi-Modal Speech In-Painting
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
von: Elyaderani, Mahsa Kadkhodaei, et al.
Veröffentlicht: (2024)
Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
von: Natesan, Sanjay, et al.
Veröffentlicht: (2024)
von: Natesan, Sanjay, et al.
Veröffentlicht: (2024)
MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage
von: Tan, Hao Hao, et al.
Veröffentlicht: (2024)
von: Tan, Hao Hao, et al.
Veröffentlicht: (2024)
Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Efficient Fine-Grained Guidance for Diffusion Model Based Symbolic Music Generation
von: Zhu, Tingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Tingyu, et al.
Veröffentlicht: (2024)
HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset
von: Saini, Shivam, et al.
Veröffentlicht: (2024)
von: Saini, Shivam, et al.
Veröffentlicht: (2024)
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
von: Shao, Keren, et al.
Veröffentlicht: (2025)
von: Shao, Keren, et al.
Veröffentlicht: (2025)
The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation
von: Nagarajan, Ashwin, et al.
Veröffentlicht: (2025)
von: Nagarajan, Ashwin, et al.
Veröffentlicht: (2025)
JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models
von: Li, Peike, et al.
Veröffentlicht: (2023)
von: Li, Peike, et al.
Veröffentlicht: (2023)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
von: Long, Phillip, et al.
Veröffentlicht: (2024)
von: Long, Phillip, et al.
Veröffentlicht: (2024)
On the de-duplication of the Lakh MIDI dataset
von: Choi, Eunjin, et al.
Veröffentlicht: (2025)
von: Choi, Eunjin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
von: Verma, Prateek
Veröffentlicht: (2023) -
Content Adaptive Front End For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2023) -
Neural Style Transfer for Audio Spectograms
von: Verma, Prateek, et al.
Veröffentlicht: (2018) -
Generative AI for Music and Audio
von: Dong, Hao-Wen
Veröffentlicht: (2024) -
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
von: Cheng, Hao, et al.
Veröffentlicht: (2025)