CAFA: a Controllable Automatic Foley Artist
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Benita, Roi, Finkelson, Michael, Halperin, Tavi, Sterkin, Gleb, Adi, Yossi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NAST: Noise Aware Speech Tokenization for Speech Language Models
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
von: Benita, Roi, et al.
Veröffentlicht: (2023)
von: Benita, Roi, et al.
Veröffentlicht: (2023)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
von: Aziz, Shiran, et al.
Veröffentlicht: (2024)
von: Aziz, Shiran, et al.
Veröffentlicht: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis
von: Wang, Junnuo
Veröffentlicht: (2025)
von: Wang, Junnuo
Veröffentlicht: (2025)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
von: Coelho, Guilherme
Veröffentlicht: (2025)
von: Coelho, Guilherme
Veröffentlicht: (2025)
Latent Watermarking of Audio Generative Models
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
von: Roman, Robin San, et al.
Veröffentlicht: (2024)
WHISTRESS: Enriching Transcriptions with Sentence Stress Detection
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
StressTest: Can YOUR Speech LM Handle the Stress?
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
Salmon: A Suite for Acoustic Language Model Evaluation
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
von: Roth, Amit, et al.
Veröffentlicht: (2024)
von: Roth, Amit, et al.
Veröffentlicht: (2024)
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
StereoFoley: Object-Aware Stereo Audio Generation from Video
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2025)
von: Tal, Or, et al.
Veröffentlicht: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
von: Chang, Xuankai, et al.
Veröffentlicht: (2024)
von: Chang, Xuankai, et al.
Veröffentlicht: (2024)
FoleyBench: A Benchmark For Video-to-Audio Models
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
Low-Resource Audio Codec (LRAC): 2025 Challenge Description
von: Wojcicki, Kamil, et al.
Veröffentlicht: (2025)
von: Wojcicki, Kamil, et al.
Veröffentlicht: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
von: Shan, Sizhe, et al.
Veröffentlicht: (2025)
von: Shan, Sizhe, et al.
Veröffentlicht: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
Automatic Music Mixing using a Generative Model of Effect Embeddings
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
Automatic Detection and Annotation of Sperm Whale Codas
von: Gubnitsky, Guy, et al.
Veröffentlicht: (2024)
von: Gubnitsky, Guy, et al.
Veröffentlicht: (2024)
Phonetic Richness for Improved Automatic Speaker Verification
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
BWSNet: Automatic Perceptual Assessment of Audio Signals
von: Veillon, Clément Le Moine, et al.
Veröffentlicht: (2023)
von: Veillon, Clément Le Moine, et al.
Veröffentlicht: (2023)
Automatic Melody Reduction via Shortest Path Finding
von: Wang, Ziyu, et al.
Veröffentlicht: (2025)
von: Wang, Ziyu, et al.
Veröffentlicht: (2025)
A Dataset for Automatic Assessment of TTS Quality in Spanish
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
Slamming: Training a Speech Language Model on One GPU in a Day
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
Exploiting Music Source Separation for Automatic Lyrics Transcription with Whisper
von: Syed, Jaza, et al.
Veröffentlicht: (2025)
von: Syed, Jaza, et al.
Veröffentlicht: (2025)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2024)
Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2024)
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
NAST: Noise Aware Speech Tokenization for Speech Language Models
von: Messica, Shoval, et al.
Veröffentlicht: (2024) -
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
von: Benita, Roi, et al.
Veröffentlicht: (2023) -
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
von: Zeldes, Ella, et al.
Veröffentlicht: (2024) -
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024) -
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
von: Aziz, Shiran, et al.
Veröffentlicht: (2024)